A venue internet of things multimedia monitoring management system

By collecting and analyzing the behavioral characteristic data of the venue audience and dynamically adjusting the interpretation model, the problem of the inability to provide personalized interpretation in existing technologies is solved, and the user experience and product sales effect are improved.

CN120337043BActive Publication Date: 2025-10-14BEIJING ASIA SATELLITE COMM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510829036.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-14
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

Existing venue guide technology is unable to provide personalized interpretation services based on the audience's real-time characteristics, resulting in low user acceptance, affecting customer experience and product sales.

Method used

The behavior perception module collects the target population's voice interaction and movement trajectory data, analyzes and matches the basic population categories, calls the corresponding interpretation model, and adjusts the speech speed, volume, language style and content depth in real time during the interpretation process to meet the needs of different groups of people.

Benefits of technology

It enables personalized adjustment of commentary content, improves user acceptance of commentary, enhances customer experience, promotes products efficiently, and reduces resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337043B_ABST
    Figure CN120337043B_ABST
Patent Text Reader

Abstract

The application discloses a kind of venue Internet of Things multimedia monitoring management systems, it is related to multimedia monitoring technical field, behavior sensing module, the behavior characteristic data of target crowd is collected, the behavior characteristic data at least includes voice interaction data, mobile trajectory data, wherein language interaction data includes preparation word and execution word;Data analysis module, the behavior characteristic data collected is analyzed and handled, extracts the key feature information that can reflect the acceptance degree of target crowd to interpretation model, the voice interaction and mobile trajectory data of customer are collected in real time by the application, the demand of different crowd can be accurately analyzed, generates the personalized interpretation model in line with current crowd, so that interpretation service is more and more accurate, is suitable for different types of crowd.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of multimedia monitoring technology, and in particular to a venue Internet of Things multimedia monitoring and management system. Background Art

[0002] Smart venues are the product of the deep integration of traditional venues and digital technology. Their core lies in intelligent perception, data-driven decision-making, and personalized service. Through technological iteration, smart venues are evolving from single-function venues to comprehensive service platforms that integrate safety, efficiency, and user experience. In the future, with the widespread adoption of AIoT technology, smart venues will play a more important role in smart cities.

[0003] In the field of venue operation and management, the Internet of Things multimedia monitoring and management system relies on advanced technologies such as sensors and network communications to efficiently complete data collection and equipment linkage. It has become a key tool for the operation and management of many venues, providing all-round data insights for venue operations.

[0004] Existing venue guide technologies, such as intelligent robots, mostly operate based on preset programs or rely on user-initiated selections, lacking real-time targeting and making it difficult to meet the personalized interpretation needs of different audiences. In contrast, the data advantages of the IoT multimedia monitoring and management system can effectively make up for this shortcoming. The system collects audience characteristic data in real time through sensors, analyzes and processes it, and transmits the information to the corresponding multimedia language device, which then calls the interpretation model that suits the current user, thereby realizing the switching and output of the corresponding interpretation content.

[0005] In this context, how to achieve accurate matching of interpretation services with the real-time characteristics of the audience based on the interpretation model switching mechanism of the Internet of Things multimedia monitoring and management system, and further enable the corresponding multimedia language equipment to provide more targeted interpretation services according to the different characteristics of the visitors, has become a difficult problem that needs to be overcome urgently. Summary of the Invention

[0006] The purpose of the present invention is to provide a venue Internet of Things multimedia monitoring and management system, which solves the technical problem of dynamically switching interpretation models for different groups of people to improve their acceptance.

[0007] A venue Internet of Things multimedia monitoring and management system, comprising:

[0008] A behavior perception module collects behavioral characteristic data of the target population, wherein the behavioral characteristic data includes at least voice interaction data and movement trajectory data, wherein the language interaction data includes preparation words and execution words;

[0009] The data analysis module analyzes and processes the collected behavioral feature data to extract key feature information that can reflect the target population's acceptance of the explanation model;

[0010] The model calling module matches the target population with the preset population classification based on the key feature information, determines the basic population category to which the target population belongs, and calls the corresponding basic interpretation model for the basic population category;

[0011] The data comparison module continuously monitors the real-time behavioral characteristic data of the target population during the interpretation process and compares it with the expected behavioral characteristic data under the invoked basic interpretation model;

[0012] The model adjustment module dynamically adjusts the current commentary model based on the type of difference from one or more dimensions including commentary speed, volume, language style, and content depth when the difference between real-time behavior feature data and expected behavior feature data exceeds a preset threshold, and generates an adapted current commentary model to improve user acceptance of the commentary content.

[0013] As a further solution of the present invention: the preset population classification includes at least the elderly, middle-aged people, and young people. The basic interpretation model corresponding to the elderly is a model with slow speaking speed and loud voice, the basic interpretation model corresponding to the middle-aged people is a model with moderate speaking speed and refined content, and the basic interpretation model corresponding to the young people is a model with fast speaking speed and lively language style.

[0014] As a further solution of the present invention: the analysis and processing of the collected behavioral characteristic data includes:

[0015] Perform speech recognition and semantic analysis on voice interaction data, extracting voice response duration, question frequency, semantic understanding accuracy, preparation words, execution words, and switching words as key feature information;

[0016] Analyze the movement trajectory data and extract the dwell time, movement speed, and movement direction change frequency as key feature information;

[0017] The operation instruction data is analyzed to extract the number of instruction repetitions, instruction execution success rate, and instruction issuance interval as key feature information.

[0018] As a further solution of the present invention: the model adjustment module further includes a preparatory switching module and a switching operation module:

[0019] The preparatory switching module, when acquiring the voice interaction data of the target population, determines the age classification result of the current target population if the preparatory word is detected. Based on the age classification result, it preloads the interpretation model that meets the current age classification from the preset database and starts the timer for countdown.

[0020] Switch the operation module. During the timing process, the key feature data of the target population is continuously obtained. If key behavioral actions or execution words are obtained, the explanation model matching the current age classification is immediately called for explanation. If no key behavioral actions or execution words are obtained after the timing ends, the prepared words are used as loading words to load the explanation model related to the loading words in the preset database.

[0021] As a further solution of the present invention: during the explanation process, the real-time behavioral characteristic data of the target group is continuously monitored and compared with the expected behavioral characteristic data under the called basic explanation model, specifically:

[0022] Establish a database of correspondence between basic interpretation models and expected behavioral characteristic data;

[0023] Collect the behavioral characteristic data of the target population in real time and match it with the expected behavioral characteristic data corresponding to the basic interpretation model called in the corresponding relationship database;

[0024] Calculate the difference between the real-time behavior feature data and the expected behavior feature data, and determine whether the difference exceeds a preset threshold.

[0025] As a further solution of the present invention: when multiple prepared words are obtained, the following steps are performed:

[0026] Determine whether the prepared words refer to the same meaning, and calculate whether the proportion of the prepared words that refer to the same meaning exceeds a preset value;

[0027] If the proportion of prepared words pointing to the same meaning is greater than the preset value, the corresponding interpretation model in the preset database is retrieved in advance for preloading;

[0028] If the proportion of prepared words pointing to the same meaning is less than or equal to the preset value, then it is continuously obtained within the preset time whether other prepared words continue to appear, and it is determined in real time within the preset time whether the newly added prepared words exceed the preset value after being added. If it exceeds, the corresponding interpretation model in the preset database is retrieved in advance for preloading. If it does not exceed, the filtered prepared words are preloaded as loading words.

[0029] As a further solution of the present invention, the screened prepared words are preloaded as loading words, specifically:

[0030] Use the basic demand words among the multiple prepared words as the main loading words, and the other words as parameter words;

[0031] The interpretation model corresponding to the main loading word is selected as the basic interpretation method, and the basic interpretation method is adjusted in combination with the corresponding interpretation model parameters in the parameter word.

[0032] As a further solution of the present invention, a derivative recommendation module is also included, and the specific workflow is as follows:

[0033] When the data analysis module obtains a prepared word from the voice interaction data, it automatically generates a set of potential demand words related to the prepared word;

[0034] Storing the potential demand word set as a pre-loaded item in a temporary cache area and starting a demand monitoring timer;

[0035] During the timing of the demand monitoring timer, continuously monitoring the voice interaction data of the target group;

[0036] If an action word matching any word in the potential demand word set is detected in the target group's voice interaction data, the model calling module retrieves the derivative product explanation model corresponding to the action word, fuses the derivative product explanation model with the current basic explanation model, and recommends it to the target group;

[0037] If no matching execution word is detected after the demand monitoring timer expires, the pre-loaded items in the temporary buffer area are cleared.

[0038] As a further solution of the present invention: during the timing of the demand monitoring timer, if a specific behavior combination is detected in the target population, the derivative recommendation is triggered faster: when the user stays in front of a product for more than a preset time and the prepared words appear in the voice interaction, the derivative product explanation model is immediately called up without waiting for the demand monitoring timer to end.

[0039] As a further solution of the present invention: it also includes a storage module, which is used to store preset crowd classifications, basic interpretation models, a database of correspondences between basic interpretation models and expected behavioral feature data, and a correspondence between prepared words and interpretation models.

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] This invention collects real-time customer voice interaction and movement trajectory data and adjusts the current explanation model based on the type of difference, taking into account one or more dimensions: commentary speed, volume, language style, and content depth. This allows for precise analysis of the needs of different groups and generates a personalized explanation model tailored to that group. This allows for continuous monitoring of customer responses throughout the explanation process, which not only enhances the customer shopping experience and makes it easier for them to understand product information, but also helps promote products efficiently and reduce ineffective resource investment. Accumulated data can also be used to optimize the explanation model, making the explanation service increasingly precise and suitable for different types of people. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 A system framework structure diagram of the present application. DETAILED DESCRIPTION

[0043] The technical solutions of the present application will be described below in conjunction with embodiments. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of the present application.

[0044] In various types of venues, the venue space is usually divided according to the actual situation, and different regions are placed with different commodity exhibits. Because different crowds have different interests in the exhibits in each region of the venue, based on industry common sense, the smart venue often uses intelligent robots, AR / VR guide systems, voice interaction devices and other technical means in the guide and explanation services, so as to reduce the dependence on manual explanation. The common use of multimedia monitoring in the venue is mainly real-time monitoring and recording of relevant evidence, but at present, multimedia monitoring is mostly an independent system, which is not deeply integrated with the guide service, and lacks analysis of the actual situation of the stayers in the venue according to their characteristics, so as to provide more targeted interpretation by the corresponding multimedia language devices according to different characteristics of the stayers.

[0045] Specifically, the existing venue guide technology (such as intelligent robots, AR / VR devices) mostly relies on preset programs or user active selection, and cannot provide personalized interpretation services according to the real-time characteristics (such as age, interest, stay duration) of the stayers. For example, when facing different users showing interest in the current commodity exhibits, the existing multimedia devices generally use a single language style to explain. For the elderly with hearing impairment, if the speed is too fast and the voice is too small, they may not be able to hear clearly, so they choose to leave; while for young people, if the speed is too slow and the voice is too loud, they may feel troublesome and leave, thereby affecting the overall experience effect of the customers.

[0046] Obviously, if personalized interpretation services cannot be provided according to the characteristics of different crowds, the stay time of the users will be greatly shortened, which will undoubtedly have an adverse effect on the sales of commodities.

[0047] Please refer to Figure 1 The present application provides a venue Internet of Things multimedia monitoring management system, which comprises:

[0048] The behavior perception module collects the behavior characteristic data of the target crowd, and the behavior characteristic data at least includes voice interaction data and movement trajectory data, wherein the voice interaction data includes preparation words and execution words.

[0049] The data analysis module analyzes and processes the collected behavioral feature data to extract key feature information that can reflect the target population's acceptance of the explanation model;

[0050] The model calling module matches the target population with the preset population classification based on the key feature information, determines the basic population category to which the target population belongs, and calls the corresponding basic interpretation model for the basic population category;

[0051] The data comparison module continuously monitors the real-time behavioral characteristic data of the target population during the interpretation process and compares it with the expected behavioral characteristic data under the invoked basic interpretation model;

[0052] The model adjustment module dynamically adjusts the current commentary model based on the type of difference from one or more dimensions including commentary speed, volume, language style, and content depth when the difference between real-time behavior feature data and expected behavior feature data exceeds a preset threshold, and generates an adapted current commentary model to improve user acceptance of the commentary content.

[0053] As an optional embodiment, during the explanation process, the real-time behavioral characteristic data of the target group is continuously monitored and compared with the expected behavioral characteristic data under the called basic explanation model, specifically:

[0054] Establish a database of correspondence between basic interpretation models and expected behavioral characteristic data;

[0055] Collect the behavioral characteristic data of the target population in real time and match it with the expected behavioral characteristic data corresponding to the basic interpretation model called in the corresponding relationship database;

[0056] Calculate the difference between the real-time behavior feature data and the expected behavior feature data, and determine whether the difference exceeds a preset threshold.

[0057] As an optional embodiment, the preset population classification includes at least the elderly, middle-aged people, and young people. The basic interpretation model corresponding to the elderly is a model with slow speaking speed and loud voice, the basic interpretation model corresponding to the middle-aged people is a model with moderate speaking speed and refined content, and the basic interpretation model corresponding to the young people is a model with fast speaking speed and lively language style.

[0058] The behavioral characteristics of the target group are tracked through cameras, Wi-Fi positioning, Bluetooth beacons and other technologies to track the target group's movement paths (movement trajectories) within the exhibition area. For example, through the camera's video analysis technology, the user's stay time in front of different product exhibits and other data can be recorded to determine the user's points of interest;

[0059] Among them, microphone arrays or intelligent voice devices are used to capture the voice information of the target population in the product display area in real time, and the collected data is pre-processed by noise reduction, format conversion, etc. The above methods are all existing technologies, and other methods can also be used to achieve them, which will not be elaborated here.

[0060] Among them, preparatory words refer to the preparatory language used by users before expressing their needs, such as "I want to know" and "this product", which are used to identify users' intention to interact; execution words are keywords that clearly express needs or feedback, such as "function", "price", "user experience", etc.

[0061] Natural language processing (NLP) technology is used to perform semantic and sentiment analysis on voice interaction data, helping us understand the core needs of users’ questions.

[0062] The mobile trajectory data is analyzed by combining the display area layout. If a user stays in front of a product for longer than a preset time, it indicates that the user is very interested in the product. The above analysis results are converted into quantitative indicators, such as the frequency of voice interaction and the duration of the mobile trajectory, as key characteristic information reflecting the degree of user acceptance.

[0063] Among them, according to the common user types in the product display scenarios, multiple basic population categories are preset, and the current visitors are divided into three types: elderly, middle-aged and young people according to behavioral characteristic data information. These divisions are not absolute and should be flexibly adjusted according to specific circumstances in actual applications.

[0064] Among them, machine learning classification algorithms (such as support vector machines and random forests) are used, key feature information is used as input, and the training model is used to achieve automatic matching of population categories.

[0065] After determining the target demographic's basic demographic category, the model call module retrieves and preloads the corresponding basic interpretation model from a preset database stored on a cloud server, reducing user wait time. This preset database can be managed using a relational database (such as MySQL), enabling fast retrieval and call access through efficient SQL queries. Multiple preset models can be pre-configured and stored in the preset database based on actual circumstances, but this will not be elaborated on here.

[0066] Among them, a corresponding relationship database is constructed through historical data accumulation and experimental testing.

[0067] Among them, by collecting the behavioral characteristic information of the target population in real time and matching it with the expected behavioral characteristic data in the corresponding relationship database, each basic interpretation model is pre-set with the expected behavioral characteristics that match it;

[0068] Furthermore, for numerical data (such as voice response duration and movement speed), the absolute difference method is used to calculate the difference value, specifically: For numerical data (such as voice response duration and movement speed), the absolute difference method is used to calculate the difference value: The comprehensiveness uses the weighted average method to calculate the overall difference between the real-time behavior feature data and the expected behavior feature data; and assigns different weights according to the importance of each feature indicator. The formula is: , n represents the number of behavioral characteristic indicators involved in calculating the overall difference, i is an index variable from 1 to n, which is used to refer to each specific behavioral characteristic indicator in turn, Indicates the importance of the i-th behavioral characteristic index in the overall difference calculation, Refers to the difference between the real-time data of the i-th behavioral characteristic indicator and the expected data under the corresponding basic interpretation model;

[0069] Furthermore, when the calculated overall difference exceeds a preset threshold, the data comparison module sends a difference warning signal to the model adjustment module, triggering the model adjustment process; if it does not exceed the threshold, the current interpretation model continues to be maintained.

[0070] After receiving the difference warning signal sent by the data comparison module, the model adjustment module first conducts an in-depth analysis of the difference data to determine the type of difference.

[0071] The core of this invention is to adjust the current commentary model based on the type of difference, taking into account one or more dimensions: commentary speed, volume, language style, and content depth. If the commentary speed is determined to be an issue, the commentary model is adjusted by modifying the speech synthesis engine's speech speed parameters. The commentary volume can also be changed by adjusting the audio playback device's volume output parameters (such as setting the volume percentage in multimedia playback software) or the speech synthesis engine's volume parameters. The current language style is replaced with a corresponding template from a preset language style template library, which includes various style templates such as professional, popular, and lively. Based on the user's comprehension ability and interest, content at different levels is extracted from the product knowledge graph and recombined. For users with weaker comprehension abilities, the system reduces the complex principle introduction and increases practical application cases.

[0072] In this way, the customer's voice interaction and movement trajectory data can be collected in real time. For example, by hearing prepared words such as "let me introduce it" or recording the time they stay in front of a certain product, the needs of different groups of people can be accurately analyzed, and a personalized explanation model that suits the current group of people can be generated. Then, customer reactions can be continuously monitored during the explanation process. For example, if it is found that customers ask questions frequently, the depth of the content can be adjusted to make the explanation more in line with their acceptance level. In this way, not only the customer's shopping experience is improved, making it easier for them to understand product information, but it can also help to promote products efficiently, reduce ineffective resource investment, and optimize the explanation model through accumulated data, making the explanation service more and more accurate and suitable for different types of people.

[0073] As an optional embodiment, analyzing and processing the collected behavioral characteristic data includes:

[0074] Perform speech recognition and semantic analysis on voice interaction data, extracting voice response duration, question frequency, semantic understanding accuracy, preparation words, execution words, and switching words as key feature information;

[0075] Analyze the movement trajectory data and extract the dwell time, movement speed, and movement direction change frequency as key feature information;

[0076] The operation instruction data is analyzed to extract the number of instruction repetitions, instruction execution success rate, and instruction issuance interval as key feature information.

[0077] This application further proposes to analyze and process the collected behavioral feature data. Specifically, a microphone array (e.g., a 4-microphone ring array) deployed in the product display area is used to collect voice signals. Then, a speech recognition engine (e.g., iFlytek offline SDK) is used to convert the analog audio into a text stream.

[0078] Then, natural language processing (NLP) technology is used to perform semantic analysis to extract voice response duration, question frequency, semantic understanding accuracy, preparation words, action words, and switching words;

[0079] For example, the time difference between the end of the system's commentary and the user's response is calculated as the audio response duration. The number of questions asked per minute is counted as the question frequency. The semantic understanding accuracy is calculated by calculating the cosine similarity between the user's response content and the standard answer in the product knowledge base.

[0080] Preparation word / execution word extraction uses regular expressions to match a preset keyword library, while switch words capture the vocabulary that indicates user intent. These are all existing technologies and will not be elaborated on in detail.

[0081] The analysis of the moving trajectory data can utilize the UWB positioning system (such as Decawave DW1000 chip) to collect real-time coordinates of the user, the positioning accuracy reaches 10 cm, and the moving trajectory is recorded at a frequency of 5 Hz, so that parameters such as the extraction of the residence time, the moving speed, and the moving direction change frequency can be obtained, which are prior art and will not be elaborated too much.

[0082] Among the display devices supporting touch interaction (such as smart shopping guide screens), the instruction repetition number, the instruction execution success rate, and the instruction issuing interval time can be obtained by capturing user operation instructions through an event listening mechanism, which are prior art and will not be elaborated too much.

[0083] As an optional embodiment, the model adjusting module further includes a preliminary switching module and a switching operation module:

[0084] The preliminary switching module, when obtaining the voice interaction data of the target crowd, determines the age classification result of the current target crowd after detecting the preparation word, preloads the explanation model in the preset database that meets the current age classification based on the age classification result, and starts a timer for countdown;

[0085] The switching operation module, during the timing process, continuously obtains the key feature data of the target crowd, and if the key behavior action or the execution word is obtained, the explanation model matched with the current age classification is immediately called for explanation; if the key behavior action or the execution word is not obtained after the timing ends, the preparation word is used as a loading word to call the explanation model related to the loading word in the preset database for loading.

[0086] It is further proposed in the application that in the commodity display scene, when the behavior perception module obtains the voice interaction data of the target crowd, the preliminary switching module first performs real-time monitoring on the data. Once the preparation word (such as “I want to know” “Give me an introduction”) is detected, the age classification process is immediately started.

[0087] First, the classification process relies on key feature information extracted by the data analysis module rather than simply dividing users by age. For example, voice response duration, word usage habits, and movement trajectory are used as division parameters. The specific division method is as follows: For example, if a user's voice response is slow, their word usage is simple and direct, and their movement speed is slow, the system will use a machine learning algorithm (such as a support vector machine) to determine that they are elderly; if the user's language expression is concise, their questions are professional, and their movement speed is moderate, they will be determined to be middle-aged; if the user's language is lively, they frequently use Internet buzzwords, and their movement trajectory shows fast browsing characteristics, they will be determined to be young. Compared with simply dividing users by age, this division method is more suitable for selecting appropriate explanation models for explanations and can better improve users' acceptance of product explanations. For example, if a user is classified as a young person in terms of age, but in reality, their voice response is slow, their word usage is simple and direct, and their movement speed is slow, it is incorrect to use the basic explanation model of young people at this time;

[0088] Secondly, after determining the classification result, the pre-switching module pre-loads the interpretation model that matches the current classification from the preset database based on the result. The preset database stores basic interpretation models designed for different classification groups, such as a model with slow speech speed and loud voice for the elderly; a model with moderate speech speed and concise content for middle-aged people; and a model with fast speech speed and lively language style for young people. The pre-loading operation stores the model data in memory in advance, shortening the response time of subsequent calls;

[0089] At the same time, the preparatory switching module starts the timer to count down. The countdown duration can be flexibly set according to actual business needs, generally between 5-15 seconds.

[0090] Finally, as the timer counts down, the switching operation module continuously acquires key feature data of the target population. This data comes from the behavior perception module and the data analysis module, including action words in voice interaction, dwell time and direction changes in movement trajectories, and operation instruction data.

[0091] If a key action (such as stopping to observe a product for a long time or repeatedly touching it) or an action word (such as "tell me more" or "demonstrate") is detected during the timing process, the switching operation module immediately calls the explanation model that matches the current age category to provide an explanation. For example, if the user is identified as a young adult and the action word "demonstrate the latest features" is detected, the basic explanation model corresponding to young people will be quickly called to introduce the product's innovative features in detail in a lively language style.

[0092] If the key action or action word is not obtained after the timer expires, the switching operation module will use the prepared word as the loading word and retrieve the explanation model related to the loading word from the preset database for loading. For example, if the prepared word is "I want to learn", the system will retrieve and load a common product introduction model from the database. This model contains basic product information, core selling points, and other content, providing users with preliminary explanation services. This helps prevent user churn while waiting, and helps ensure that the system can continuously and accurately provide personalized product explanation services to the target audience, improving user acceptance of the explanation content and shopping experience.

[0093] As an optional embodiment, when multiple prepared words are obtained, the following steps are performed:

[0094] Determine whether the prepared words refer to the same meaning, and calculate whether the proportion of the prepared words that refer to the same meaning exceeds a preset value;

[0095] If the proportion of prepared words pointing to the same meaning is greater than the preset value, the corresponding interpretation model in the preset database is retrieved in advance for preloading;

[0096] If the proportion of prepared words pointing to the same meaning is less than or equal to the preset value, it is continuously obtained within the preset time whether other prepared words continue to appear, and it is determined in real time within the preset time whether the newly added prepared words exceed the preset value after being added. If it exceeds, the corresponding interpretation model in the preset database is retrieved in advance for preloading. If it does not exceed, the filtered prepared words are preloaded as loading words.

[0097] This application further proposes that, after the behavior perception module obtains the voice interaction data of the target population, if multiple prepared words (such as "I want to know" and "Introduce me") are detected, the word vector model (such as Word2Vec and BERT) in natural language processing (NLP) technology is first used to convert each prepared word into a vector representation in a high-dimensional space, and the distance between them in the semantic space is calculated using the cosine similarity algorithm to determine whether they point to the same meaning;

[0098] Prepared words with a similarity higher than the threshold are judged to point to the same meaning. Then, the ratio of the number of prepared words pointing to the same meaning to the total number of prepared words is calculated; if the ratio of the number of prepared words pointing to the same meaning exceeds the preset value, it means that the user needs are clear and concentrated. At this time, the system immediately executes the operation of the original preparation switching module: according to the age classification results of the target population (determined by analyzing voice intonation, word usage habits, movement trajectory and other features), the corresponding explanation model that meets the age classification in the preset database is retrieved in advance for preloading. For example, if the user is determined to be a middle-aged person and the prepared words all point to "understanding the price-performance ratio of the product", the refined explanation model corresponding to the middle-aged person is preloaded, focusing on price, performance and other information;

[0099] If the percentage does not exceed the preset value, the system activates a time window monitoring mechanism, continuously monitoring for the appearance of new prepared words within the preset time. During this period, with each new prepared word, the system recalculates the percentage of prepared words with the same meaning in real time and determines whether it exceeds the preset value. For example, if the initial three prepared words only account for 60%, but two more prepared words with the same meaning are added within 8 seconds, bringing the percentage to 80%, the preloading process is immediately triggered.

[0100] If the percentage does not exceed the preset value after the preset time, it means that the user demand is relatively scattered or vague. At this time, the prepared words are screened and integrated, and the common explanation models related to these loaded words in the preset database are retrieved for pre-loading, briefly covering all aspects of the basic information of the product, and providing users with a preliminary explanation;

[0101] Intelligent filtering logic has been added to the preload decision-making process after detecting the prepared word. This allows for quick responses when user needs are clear and flexible responses when they are ambiguous. This helps avoid model mismatches caused by misjudged needs, significantly improving user experience and system response efficiency.

[0102] As an optional embodiment, the filtered prepared words are preloaded as loading words as follows:

[0103] Use the basic demand words among the multiple prepared words as the main loading words, and the other words as parameter words;

[0104] The interpretation model corresponding to the main loading word is selected as the basic interpretation method, and the basic interpretation method is adjusted in combination with the corresponding interpretation model parameters in the parameter word.

[0105] Specifically, the filtered prepared words are preloaded as loading words. The basic demand words among the prepared words are used as the main loading words, and the other words are used as parameter words. By preloading key explanation resources into memory or related modules, it can quickly respond to user interactions. The main loading words represent the core needs of users, while the parameter words provide additional personalized information.

[0106] Secondly, select the explanation model corresponding to the main loading words as the basic explanation method. The basic explanation model is pre-set by the system to meet the basic needs of most users. For example, in the explanation system, the basic explanation model may be a general explanation template used to introduce the basic information of the exhibits, such as the name, material, and purpose of the exhibits;

[0107] The basic interpretation mode is adjusted in combination with the corresponding explanation model parameters in the parameter word. The information in the parameter word (such as language style, speech speed, volume, etc.) is used to personalize the adjustment of the basic interpretation model. For example, if the parameter word contains "elderly people", the system may adjust the interpretation model to have a slower speech speed and moderate volume to meet the hearing needs of the elderly; if the parameter word contains "young people", the interpretation model may be adjusted to have a faster speech speed and normal volume to meet the preferences of young people.

[0108] As an optional embodiment, a derivative recommendation module is further included, and the specific workflow is as follows:

[0109] When the data analysis module obtains the preparation word from the voice interaction data, a set of potential demand words related to the preparation word is automatically generated;

[0110] The set of potential demand words is stored in a temporary cache area as a preloaded item, and a demand monitoring timer is started;

[0111] During the timing of the demand monitoring timer, the voice interaction data of the target population is continuously monitored;

[0112] If a matching execution word is detected in the voice interaction data of the target population, the derivative product interpretation model corresponding to the execution word is called through the model calling module, and the derivative product interpretation model is fused with the current basic interpretation model to recommend to the target population;

[0113] If no matching execution word is detected when the demand monitoring timer expires, the preloaded items in the temporary cache area are emptied.

[0114] Specifically, when the data analysis module identifies the preparation word (such as "this mobile phone" or "this furniture") from the voice interaction data, a semantic expansion algorithm is immediately started. Based on the product knowledge graph and natural language processing technology, the algorithm generates a set of potential demand words;

[0115] First, the core functions of the product are extended, second, associated products are recommended according to the use scenarios of the product, and finally, hot product comparison words are automatically supplemented for popular products. The set of potential demand words generated by the TF-IDF algorithm is sorted by weight, and the 5-10 most relevant words are selected;

[0116] After the set of potential demand words is generated, the system stores it in a memory-type temporary cache area (such as Redis cache), and starts a demand monitoring timer (the default setting is 15 seconds). The cache area uses the LRU (Least Recently Used) eviction policy to ensure efficient data storage and fast retrieval;

[0117] During the timing process, the behavior perception module collects voice interaction data at a frequency of 10 times per second, and compares the user voice with the potential demand words in the cache area through a keyword real-time matching algorithm (such as an AC automatic machine algorithm). For example, when the user says “can you filter formaldehyde”, the system quickly matches “purification effect” in the potential demand word set, triggering the subsequent recommendation logic;

[0118] Once the matching execution word is detected, the model calling module immediately retrieves the corresponding explanation model from the derivative product preset database. The preset database pre-designs differentiated explanation strategies for different commodity categories and demand scenarios, and the retrieved derivative product explanation model is fused with the current basic explanation model (main model based on user age classification) through a weight allocation algorithm. For example, if the user is young, the basic explanation model accounts for 60% (maintaining a lively language style), and the derivative model accounts for 40% (highlighting professional parameters), and finally a hybrid explanation content of “young expression + deep function analysis” is formed. The system converts the fused content into fluent voice through a voice synthesis engine (such as a Keda Xunfei multi-style synthesis), and presents the text information on the display screen;

[0119] If the demand monitoring timer ends without detecting a matching execution word, the system automatically clears the temporary cache area and releases memory resources. When the preloaded basic explanation model has not been triggered, the monitoring results of the derivative recommendation module can be used as supplementary basis for adjusting the preloaded model. For example, if “cost performance” appears frequently in the potential demand word set, the priority of the middle-aged basic explanation model can be adjusted in advance. The un-matched potential demand word data will be fed back to the data analysis module for optimizing the subsequent potential demand word generation algorithm;

[0120] Through the above process, the derivative recommendation module realizes the deep mining of explicit demand to implicit demand of the user, improves the accuracy of commodity recommendation under the premise of not interrupting the normal explanation process with light-weight monitoring and intelligent fusion strategy, avoids user resistance caused by excessive recommendation, and finally enhances the user's acquisition efficiency and purchase willingness of commodity information.

[0121] As an optional embodiment, during the timing process of the demand monitoring timer, if a specific behavior combination of the target population is detected, the derivative recommendation is accelerated: when the user stays in front of a certain commodity for more than a preset time, and the preparation word appears in the voice interaction, the derivative product explanation model is immediately retrieved without waiting for the demand monitoring timer to end.

[0122] Specifically, during the operation of the demand monitoring timer, the system simultaneously monitors the target population's movement trajectory data and voice interaction data, forming dual trigger conditions. The UWB positioning system tracks the user's location in real time. When the user's stay time in a certain product display area exceeds the preset threshold, the behavior is marked as a person of deep concern and the microphone array simultaneously captures the voice data. Once the preparation word is recognized, the subsequent process is triggered. Only when the above two conditions are met at the same time will the demand monitoring timer waiting stage be skipped and the derivative recommendation will be started directly.

[0123] Secondly, by obtaining the triggering combination of specific behaviors, the system immediately retrieves the potential demand word set related to the prepared word (using the original generation logic) and quickly extracts the pre-loaded derivative product explanation model from the temporary cache. For example, if a user stays in the mobile phone exhibition area for 3 minutes and asks "this phone", the system will instantly load the explanation content corresponding to potential needs such as "performance parameters" and "battery life";

[0124] Compared to regular timed triggers, derivative recommendations triggered by specific behavioral combinations receive higher priority. The system will temporarily increase the weight of the derivative model in the fusion strategy to prioritize core information that users are likely to care about, such as explaining the phone's processor performance and battery capacity. At the same time, the language style characteristics of the basic explanation model are retained. By cross-validating dwell time and voice commands, this effectively filters out casual browsing behavior, helps focus on users with real needs, and avoids ineffective recommendations.

[0125] As an optional embodiment, a storage module is further included, which is used to store preset crowd classifications, basic interpretation models, a database of correspondences between basic interpretation models and expected behavioral feature data, and a database of correspondences between prepared words and interpretation models.

[0126] Specifically, the storage module adopts a distributed storage architecture, combining relational databases (such as MySQL) and unstructured databases (such as MongoDB) to implement data classification management. Data is stored through the storage module to facilitate subsequent calls. This is an existing mature technology and will not be elaborated on here.

[0127] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A venue Internet of Things multimedia monitoring and management system, characterized in that: include: The behavior perception module collects behavioral characteristic data of the target population, including at least voice interaction data and movement trajectory data. The voice interaction data includes preparatory words and action words. Preparatory words refer to the preparatory language used by users before expressing their needs, which are used to identify users' intention to interact. Action words are keywords that clearly express needs or feedback. The data analysis module analyzes and processes the collected behavioral feature data to extract key feature information that can reflect the target population's acceptance of the explanation model; The model calling module matches the target population with the preset population classification based on the key feature information, determines the basic population category to which the target population belongs, and calls the corresponding basic interpretation model for the basic population category; The data comparison module continuously monitors the real-time behavioral characteristic data of the target population during the interpretation process and compares it with the expected behavioral characteristic data under the invoked basic interpretation model; The model adjustment module dynamically adjusts the current commentary model based on the type of difference from one or more dimensions including commentary speed, volume, language style, and content depth when the difference between real-time behavior feature data and expected behavior feature data exceeds a preset threshold, and generates an adapted current commentary model to improve user acceptance of the commentary content.

2. The venue Internet of Things multimedia monitoring and management system according to claim 1 is characterized in that: The preset population classification includes at least the elderly, middle-aged people, and young people. The basic interpretation model corresponding to the elderly is a model with slow speaking speed and loud voice, the basic interpretation model corresponding to the middle-aged people is a model with moderate speaking speed and refined content, and the basic interpretation model corresponding to the young people is a model with fast speaking speed and lively language style.

3. The venue Internet of Things multimedia monitoring and management system according to claim 1 is characterized in that: The analyzing and processing of the collected behavioral characteristic data includes: Perform speech recognition and semantic analysis on voice interaction data, extracting voice response duration, question frequency, semantic understanding accuracy, preparation words, execution words, and switching words as key feature information; Analyze the movement trajectory data and extract the dwell time, movement speed, and movement direction change frequency as key feature information; The operation instruction data is analyzed to extract the number of instruction repetitions, instruction execution success rate, and instruction issuance interval as key feature information.

4. The venue Internet of Things multimedia monitoring and management system according to claim 3 is characterized in that: The model adjustment module also includes a preparatory switching module and a switching operation module: The preparatory switching module, when acquiring the voice interaction data of the target population, determines the age classification result of the current target population if the preparatory word is detected. Based on the age classification result, it preloads the interpretation model that meets the current age classification from the preset database and starts the timer for countdown. Switch the operation module. During the timing process, the key feature data of the target population is continuously obtained. If key behavioral actions or execution words are obtained, the explanation model matching the current age classification is immediately called for explanation. If no key behavioral actions or execution words are obtained after the timing ends, the prepared words are used as loading words to load the explanation model related to the loading words in the preset database.

5. The venue Internet of Things multimedia monitoring and management system according to claim 1 is characterized in that: During the explanation process, the real-time behavioral characteristic data of the target group is continuously monitored and compared with the expected behavioral characteristic data under the called basic explanation model, specifically: Establish a database of correspondence between basic interpretation models and expected behavioral characteristic data; Collect the behavioral characteristic data of the target population in real time and match it with the expected behavioral characteristic data corresponding to the basic interpretation model called in the corresponding relationship database; Calculate the difference between the real-time behavior feature data and the expected behavior feature data, and determine whether the difference exceeds a preset threshold.

6. The venue Internet of Things multimedia monitoring and management system according to claim 1 is characterized in that: When multiple prepared words are obtained, perform the following steps: Determine whether the prepared words refer to the same meaning, and calculate whether the proportion of the prepared words that refer to the same meaning exceeds a preset value; If the proportion of prepared words pointing to the same meaning is greater than the preset value, the corresponding interpretation model in the preset database is retrieved in advance for preloading; If the proportion of prepared words pointing to the same meaning is less than or equal to the preset value, it is continuously obtained within the preset time whether other prepared words continue to appear, and it is determined in real time within the preset time whether the newly added prepared words exceed the preset value after being added. If it exceeds, the corresponding interpretation model in the preset database is retrieved in advance for preloading. If it does not exceed, the filtered prepared words are preloaded as loading words.

7. The venue Internet of Things multimedia monitoring and management system according to claim 6 is characterized in that: The specific steps of preloading the filtered prepared words as loading words are as follows: Use the basic demand words among the multiple prepared words as the main loading words, and the other words as parameter words; The interpretation model corresponding to the main loading word is selected as the basic interpretation method, and the basic interpretation method is adjusted in combination with the corresponding interpretation model parameters in the parameter word.

8. The venue Internet of Things multimedia monitoring and management system according to claim 1 is characterized in that: It also includes a derivative recommendation module, the specific workflow is as follows: When the data analysis module obtains a prepared word from the voice interaction data, it automatically generates a set of potential demand words related to the prepared word; Storing the potential demand word set as a pre-loaded item in a temporary cache area and starting a demand monitoring timer; During the timing of the demand monitoring timer, continuously monitoring the voice interaction data of the target group; If an action word matching any word in the potential demand word set is detected in the target group's voice interaction data, the model calling module retrieves the derivative product explanation model corresponding to the action word, fuses the derivative product explanation model with the current basic explanation model, and recommends it to the target group; If no matching execution word is detected after the demand monitoring timer expires, the pre-loaded items in the temporary buffer area are cleared.

9. The venue Internet of Things multimedia monitoring and management system according to claim 8, characterized in that: During the timing of the demand monitoring timer, if a specific behavior combination is detected in the target population, the derivative recommendation will be triggered faster: when the user stays in front of a product for more than the preset time and the prepared words appear in the voice interaction, the derivative product explanation model will be immediately called up without waiting for the demand monitoring timer to end.

10. The venue Internet of Things multimedia monitoring and management system according to claim 1, characterized in that: It also includes a storage module, which is used to store preset crowd classifications, basic interpretation models, a database of correspondences between basic interpretation models and expected behavioral feature data, and a database of correspondences between prepared words and interpretation models.

Citation Information

Patent Citations

  • Exhibition hall application method and system based on artificial intelligence

    CN117473162A

  • Intelligent interactive voice explanation system based on AIGC large model

    CN119169994A