Information processing system
Patent Information
- Application Number
- CN202610281379.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-19
- Filing Date
- 2026-03-09
- Publication Date
- 2026-09-22
AI Technical Summary
[0003]本发明要解决的课题在于:现有针对校园欺凌及相关心理风险的监测与干预手段主要依赖人工观察、事后问卷调查或被动投诉举报,具有发现时机滞后、依赖主观判断强、信息不完整以及难以大范围、长时间持续监控等问题
在一实施例中,端端安装于学校走廊,持续采集学生活动画面。服务器通过图像预处理与人脸检测模块,识别某对象在多个时间段被其他对象围在角落,服务器通过姿态估计与行为特征模块检测到多次推搡动作,并通过表情识别模块检测到该对象长时间处于恐惧与悲伤情绪。服务器在风险指标计算模块中,将这些特征综合输入到模型,得出高风险指标值。服务器随后构造提示语句,调用生成式人工智能模型生成风险说明和应对方案,并将警告信息推送给端端。用户在端端查看警告后前往现场确认,若确认存在攻击性行为,则在端端界面选择“确认为攻击性行为”,该评价信息反向传输至服务器,用于更新模型。
Smart Images

Figure CN122799484A_ABST
Abstract
Description
Technical Field
[0001] The technology disclosed herein relates to an information processing system. Background Technology
[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot speech in response to the user's speech.
[0003] The problem this invention aims to solve is that existing monitoring and intervention methods for school bullying and related psychological risks mainly rely on manual observation, post-incident questionnaires, or passive complaints and reports. These methods suffer from problems such as delayed detection, strong reliance on subjective judgment, incomplete information, and difficulty in large-scale, long-term continuous monitoring. Specifically, on the one hand, children who are bullied or in high-pressure interpersonal environments often find it difficult to proactively seek help from guardians or education personnel in a timely manner, resulting in bullying behavior remaining undetected and unstoppable for extended periods. On the other hand, existing methods based on simple video surveillance are labor-intensive and struggle to accurately identify segments containing potential bullying risks from large amounts of video data. Furthermore, even if some systems can initially identify emotions or behaviors, they often lack the ability to translate the identification results into actionable intervention suggestions, failing to provide guardians and education personnel with targeted coping strategies and decision support. Simultaneously, existing technologies lack mechanisms for self-learning and closed-loop optimization of relevant models using newly collected data, making it difficult to continuously improve the detection accuracy of bullying potential over time. In summary, the present invention aims to provide a system capable of automatically analyzing children's facial expressions and behaviors, detecting the possibility of bullying in real time or near real time, automatically generating and adjusting warning messages, providing specific coping strategies to guardians and education personnel, and possessing self-learning capabilities, so as to improve the timeliness, accuracy, and interventionability of bullying detection. Summary of the Invention
[0004] To address the aforementioned issues, this invention proposes an information processing system comprising a processor configured to: analyze children's facial expressions and behaviors using image recognition and emotion analysis technologies to detect the possibility of bullying; generate warning messages using a generative artificial intelligence model and adjust the warning content accordingly; and provide specific countermeasures based on the analysis results to guardians and education-related personnel. Specifically, the processor analyzes images or video frames uploaded by camera devices or terminals to automatically identify children's facial expression features and body behavior characteristics, determining whether negative emotions such as fear or sadness exist, and whether others are engaging in aggressive or bullying-related behaviors such as pushing, surrounding, or blaming the target child. Upon detecting a bullying possibility that meets preset conditions, the processor automatically triggers the generation of a warning message. The processor, through a generative artificial intelligence model, automatically generates corresponding warning content based on the detected emotional state, behavior type, context, and severity, and dynamically adjusts the warning's wording, detail, and sensitivity level according to different user roles (such as guardians, teachers, and school administrators) to improve the effectiveness and comprehensibility of information delivery. The processor is further configured to provide guardians and education personnel with a dashboard displaying the results of children's facial expression and behavior analysis. Through an accessible graphical interface, users can intuitively view emotional trends, behavioral event records, and triggered alarm history over a specific time period, facilitating a comprehensive assessment of the child's psychological and interpersonal environment. Furthermore, the processor learns from newly acquired data, periodically evaluating and updating models used for facial expression recognition, behavior recognition, and bullying probability assessment. This forms a feedback loop encompassing user feedback and actual event results, continuously optimizing model parameters and improving the accuracy and robustness of bullying probability detection. Through this structure and processing flow, the system of this invention can achieve automated early detection of child bullying risks, intelligent alerts, and actionable countermeasures without significantly increasing manpower, effectively solving problems such as delayed bullying monitoring, inaccurate identification, and lack of closed-loop learning mechanisms in existing technologies.
[0005] "System" refers to a collection of devices consisting of one or more hardware and software components, used to perform functions such as image acquisition, data processing, analysis, alarm generation, and information presentation. It includes at least a processor and a memory, network interface, and (optionally) display or interactive device that are communicatively connected to the processor.
[0006] A processor is an electronic circuit unit that executes computer program instructions to perform processing tasks such as image recognition, emotion analysis, generative artificial intelligence reasoning, data analysis, interface control, and model updating. It can be a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a programmable logic device (FPGA), or any combination thereof.
[0007] Image recognition technology refers to the technology of identifying objects, people, postures, behaviors, and other information contained in static images or video frames by extracting features and recognizing patterns. It includes, but is not limited to, deep learning-based methods for object detection, posture estimation, and action recognition.
[0008] "Emotion analysis technology" refers to the analysis of children's facial expressions, postures, and behavioral characteristics to determine their emotional state (such as happiness, sadness, anger, fear, neutrality, etc.) and its changing trends. It typically includes facial expression classification models and emotion state assessment algorithms based on machine learning or deep learning.
[0009] "Children" refers to users who are minors, including students receiving education in educational institutions such as kindergartens, primary schools, and middle schools, as well as minors under guardianship in a family environment.
[0010] "Facial expressions and behaviors" refer to the outward expressions of children over a certain period of time, such as changes in facial muscles, eye contact, mouth shape, body posture, limb movements, and ways of interacting with others, which can be recorded by images or videos and analyzed by algorithms.
[0011] "Possibility of bullying" refers to a risk assessment result obtained by the system based on image recognition and emotion analysis results, which comprehensively evaluates whether or not aggressive, exclusionary, or persistent unfriendly behavior against children exists or is occurring in the current scene. This result indicates that there is a certain probability of bullying behavior, but it is not necessarily equivalent to a final determination in a legal or disciplinary sense.
[0012] "Generative AI models" refer to AI models that learn language or multimodal information distributions through a large amount of training data, thereby automatically generating text content. For example, natural language generation models are used to generate or rewrite warning messages and countermeasure suggestions based on the analysis results of the input, scene information and user roles.
[0013] "Warning messages" are text or multimedia messages that are automatically generated by the system and sent to guardians and / or education-related personnel when the possibility of bullying is detected to reach or exceed a preset threshold. These messages are used to alert to potential risks, explain the situation, and guide subsequent intervention.
[0014] "Warning content" refers to the specific information contained in the warning message, including but not limited to the time and place of the incident, the children involved, the types of emotions and behaviors detected, the overall risk level, and corresponding suggestions or precautions.
[0015] "Guardian" refers to a natural person or organization that has the responsibility of guardianship for a child in accordance with the law or contract, including but not limited to parents, legal guardians, stepparents, grandparents or other entities that actually assume the responsibility of caring for the child.
[0016] "Education-related personnel" refers to personnel who have a working relationship with children in the process of education, teaching or student management, including but not limited to teachers, homeroom teachers, school counselors, school administrators and other staff of educational institutions.
[0017] "Specific countermeasures" refer to the actionable suggestions and plans generated by the system for guardians and education-related personnel based on the detected emotional state, behavior type, and risk level. These include communication suggestions, observation points, school intervention measures, and psychological counseling suggestions.
[0018] A "dashboard" is a visual information interface generated by the system and presented to guardians and education-related personnel. It is a graphical or mixed text and image display page used to centrally display data such as the results of children's facial expression and behavior analysis, emotional trends, event records, and alarm history.
[0019] "Interface" refers to the human-computer interaction interface used by guardians and education-related personnel to access and view information provided by the system and to perform interactive operations, including but not limited to web page interfaces, mobile application interfaces, desktop application interfaces, or other graphical user interfaces.
[0020] "Newly acquired data" refers to the latest images, video frames, analysis results, and related annotations or user feedback information that the system continuously acquires from camera equipment, terminals, or other input sources during system deployment and operation, and that have not yet been used for current model training or evaluation.
[0021] "Self-learning" refers to the process by which a system automatically or semi-automatically retrains, fine-tunes, or updates the parameters of models used for image recognition, sentiment analysis, and bullying probability determination using newly collected data and user feedback, in order to gradually improve model performance.
[0022] "Model evaluation" refers to the process by which the system evaluates the performance parameters such as accuracy, recall, and robustness of the currently used image recognition model, sentiment analysis model, and bullying detection model using preset evaluation indicators and test data, in order to determine whether the model needs to be updated or optimized.
[0023] A "feedback loop" refers to the structure and process by which the system continuously feeds back detection results, user feedback, and post-event verification information to the model training and evaluation process during operation, thus forming a closed-loop mechanism from detection—feedback—learning—update and back to detection, in order to continuously improve the accuracy of bullying probability detection and the overall system performance. Attached Figure Description
[0024] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.
[0025] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.
[0026] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.
[0027] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.
[0028] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.
[0029] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.
[0030] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.
[0031] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.
[0032] Figure 9 This represents an emotion map that maps multiple emotions.
[0033] Figure 10 This represents an emotion map that maps multiple emotions.
[0034] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.
[0035] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.
[0036] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.
[0037] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation
[0038] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.
[0039] First, let me explain the terminology used in the following instructions.
[0040] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0041] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.
[0042] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.
[0043] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.
[0044] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.
[0045] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.
[0046] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.
[0047] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0048] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.
[0049] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.
[0050] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0051] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.
[0052] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.
[0053] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0054] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0055] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.
[0056] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.
[0057] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."
[0058] In existing technologies, most solutions for detecting the risk of child bullying in school or home environments rely solely on simple image recognition or keyword rules for alerts, presenting the following technical problems: First, computer systems typically perform static analysis on single-frame images or short clips, lacking continuous modeling of children's emotional states and behavioral patterns over time. This makes it difficult to detect children who are chronically stressed or fearful, resulting in insufficient temporal recognition capabilities and judgment accuracy. Second, traditional bullying detection systems often employ fixed threshold rules and static models, exhibiting poor adaptability to complex and changing real-world scenarios. They struggle to automatically adjust model parameters and judgment criteria based on false alarms and missed alarms during actual use, lacking effective feedback learning mechanisms. Consequently, the computer processing flow cannot self-optimize as the environment and data distribution change. Third, existing early warning systems often generate alert content using pre-written fixed templates or simple text splicing, making it difficult to dynamically adjust the expression and level of detail based on different scenarios, risk levels, and recipient attributes. This leads to a mismatch between information presentation and actual user needs, reducing the effectiveness of computational results in supporting user decision-making and intervention. Fourth, in traditional technologies, the server side often designs the image analysis module and the user interaction module separately, lacking a unified data loop. This makes it impossible to effectively feed back the user's actual response results and evaluation information to the model training and evaluation process, resulting in the entire computer system failing to form a continuous improvement closed-loop optimization process at both the algorithm level and the human-computer interaction level.
[0059] In summary, existing technologies have not yet provided a system capable of providing a unified, server-side management system for the entire computer processing flow, from image acquisition, temporal emotion and behavior fusion analysis, and quantitative assessment of bullying probability, to the automatic generation of personalized warning messages and response suggestions based on generative artificial intelligence models, and continuous self-learning and updating of the identification and evaluation models based on user feedback. Therefore, it is necessary to propose a computer implementation scheme that improves server architecture, data processing flow, and model linkage mechanisms to enhance the accuracy, interpretability, and guidance capabilities for user intervention in bullying detection, thereby substantially improving the processing performance and resource utilization efficiency of computer technology in such application scenarios.
[0060] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.
[0061] In this invention, the server includes: a processing unit for receiving image information containing children's facial expressions and behaviors from an imaging device and performing decoding, frame sampling, and preprocessing; an analysis unit for extracting facial and body regions from the image information using image processing technology, and performing emotional state estimation on the facial regions, posture estimation, and action classification on the body regions to obtain temporal data of aggressive and victimized behaviors; an evaluation unit for performing feature fusion of temporal change information of emotional states with the temporal information of behaviors, calculating an evaluation value representing the probability of bullying, and outputting a bullying judgment result according to predetermined standards; a generation unit for automatically constructing prompt statements for a generative artificial intelligence model based on the judgment result, contextual information, and object information, calling the generative artificial intelligence model to generate warning messages and response policy texts that match the current scenario, and dynamically adjusting the expression mode and level of detail of the warning content according to the evaluation value and recipient attributes; and a learning unit for structurally storing the judgment result, emotion and behavior analysis summary, warning message, and response policy text, and combining the response results and evaluation information input by guardians and education-related personnel as learning data to periodically update the parameters of the emotion recognition model, action classification model, and bullying evaluation model. This allows for the construction of a closed-loop computer processing architecture on the server side, encompassing multimodal temporal feature extraction, risk quantification assessment, natural language alarm generation, and user feedback-driven adaptive learning. This enables high-precision and scalable detection of the possibility of child bullying and significantly improves the robustness, adaptability, and support for user intervention decisions in complex real-world scenarios, thereby fundamentally improving the performance of computer technology in related fields.
[0062] "Image information" refers to visual data acquired by an imaging device and represented in digital form, containing the appearance, position, and changes of objects in a scene, including single-frame images and continuous video frames.
[0063] "Imaging device" refers to a physical device used to capture image information containing children's facial expressions and behaviors, including cameras, webcams, built-in cameras in terminals, and other devices with imaging capabilities.
[0064] "Time information" refers to the time stamp associated with the acquired image information, used to indicate the specific time point or time period of image acquisition, including timestamps, time sequence numbers, etc.
[0065] "Identification information" refers to the identifying data used to distinguish different image sources or objects, including camera location identifiers, device identifiers, and anonymous object numbers, which are used to associate and manage the data in subsequent processing.
[0066] "Information processing device" refers to computing equipment that performs data processing tasks such as image decoding, feature extraction, emotion recognition, behavior analysis, evaluation calculation, and result output, including servers, computer systems, or terminals with computing capabilities.
[0067] "Secure communication methods" refer to communication technologies and protocols that ensure the confidentiality and integrity of data through encryption and authentication mechanisms during the transmission of image information, including encryption protocols based on secure transport layers and network transmission methods with access control mechanisms.
[0068] "Compression and encoding" refers to the process of reducing the data volume and converting the format of raw image information, including using video compression algorithms or encoding standards to convert raw image data into bitstreams or compressed files for transmission and storage.
[0069] "Decoding" refers to the process of restoring compressed and encoded image information to its original form as image frame data that can be used by image processing algorithms.
[0070] Image processing technology refers to a set of algorithms that perform analysis and transformation on digital images to achieve operations such as target detection, region segmentation, and feature extraction, so as to facilitate subsequent recognition and evaluation.
[0071] "Facial region" refers to a sub-region of an image containing facial features detected and located using image processing techniques. It is used for processing such as emotion state estimation and identity association.
[0072] "Body region" refers to a sub-region of an image that contains the outline of a person's body or the location of joints, detected and located using image processing technology, and is used for pose estimation and motion analysis.
[0073] "Emotional state estimation" refers to the process of extracting features from facial regions and inputting them into a recognition model to determine the type of emotion and its probability that a person exhibits at a specific moment.
[0074] "Posture estimation" refers to the process of analyzing body regions to infer the spatial position and relative relationship of various joints or body parts in order to represent the current body posture.
[0075] "Motion classification" refers to the process of identifying and labeling the types of actions performed by a person based on changes in posture or body region features over a continuous period of time, including determining whether there is aggressive or passive behavior.
[0076] "Aggressive behavior" refers to behavioral patterns that are identified through action classification as exerting physical or psychological pressure on others, including pushing, hitting, surrounding, and threatening physical actions.
[0077] "Aggressive behavior" refers to the passive behavioral patterns exhibited by the party identified as the victim of aggressive behavior through action classification and spatiotemporal relationship analysis, including being pushed, surrounded, or forcibly pulled.
[0078] "Time-varying information" refers to data on changes in emotional states or behavioral characteristics recorded over a continuous time dimension, used to reflect the trend and persistence of emotions or actions over a period of time.
[0079] "Time series information" refers to a collection of emotional states, posture estimation results, or action classification results organized in chronological order, used for time series modeling and comprehensive evaluation.
[0080] "Feature quantity" refers to a numerical representation that can be used by the evaluation model, generated by feature extraction and fusion algorithms from information on temporal changes in emotional state, behavioral time series information, and other relevant information.
[0081] "Assessment value" refers to a numerical indicator calculated by an assessment model based on characteristic quantities, used to quantify the likelihood or risk of bullying.
[0082] "Preset criteria" refers to the rules or thresholds that are pre-set in the system to interpret evaluation values and determine the likelihood of bullying, including single thresholds, multi-level thresholds, and combined conditions.
[0083] "Bullying probability assessment result" refers to the classification or grading conclusion made based on the assessment value and predetermined standards to determine whether there is a risk or tendency of bullying in the current scenario.
[0084] "Contextual information" refers to auxiliary information related to the image acquisition scene, including location type, time period, activity category, number of participants, etc., which is used to enrich the context of the generated content.
[0085] "Object information" refers to attribute information related to the object being analyzed, including the child's anonymous ID, role type, and historical risk records, which are used to distinguish and refer to the object in the prompt statements and generated content.
[0086] "Generative AI models" refer to AI models that can automatically generate natural language text, suggested content, or other data outputs based on input text or structured information, including large-scale language models or multimodal generative models based on deep learning.
[0087] "Prompt statements" refer to the input text used to provide task descriptions, scene descriptions, and output requirements to generative artificial intelligence models, guiding the model to generate warning messages and response strategies relevant to the current scene.
[0088] "Warning messages" refer to notification texts or information content that are based on the results of bullying probability assessments and the output of generative artificial intelligence models, and are used to remind guardians or education-related personnel of potential risks.
[0089] "Response Guidelines Text" refers to text content generated by generative artificial intelligence models that provides processing suggestions, communication strategies, and intervention steps for detected bullying possibilities.
[0090] "Recipient attributes" refer to the characteristic information related to the user who receives the warning message and response policy text, including user role, professional background, permission level, language preference, etc., which are used to adjust the message expression and level of detail.
[0091] "Adjustment of content and level of detail" refers to the process of dynamically adjusting the wording intensity, technical detail depth, and information content of warning messages and response policy texts based on assessment values and recipient attributes.
[0092] "Information display device" refers to a display terminal used to receive and present warning messages, response policy texts, and related analysis results, including mobile terminals, computer terminals, tablet devices, or other devices with display and interactive functions.
[0093] "Visual presentation" refers to the display method that presents the judgment results and analysis information on the display interface in the form of graphics, charts, timelines, text lists or combinations, so that users can understand them intuitively.
[0094] "Suggested content" refers to the textual information output by the generative artificial intelligence model based on prompts, which guides users to take specific actions. This includes communication suggestions, intervention plans, and precautions.
[0095] "Response Outcomes" refers to the actual actions taken by guardians or education-related personnel after receiving warning messages and response guidelines, and the records of their effects.
[0096] "Evaluation information" refers to subjective or objective feedback data from users regarding the accuracy of warnings, the applicability of suggestions, and the effectiveness of handling, which reflects the degree of consistency between system output and actual situation.
[0097] "Learning data" refers to a collection of data, consisting of original features, judgment results, user feedback, etc., used to train or update recognition models and evaluate models.
[0098] "Machine learning algorithms" refer to computational methods that use data to automatically adjust model parameters to improve prediction or classification performance, including supervised learning, semi-supervised learning, reinforcement learning, and combinations thereof.
[0099] A “recognition model” refers to a statistical or neural network model used to perform tasks such as emotion state estimation and action classification, which outputs the corresponding emotion category or action category based on input features.
[0100] An “evaluation model” refers to a mathematical or learning model used to calculate an evaluation value based on feature quantities and to determine the probability of bullying, including regression models, classification models, or rating models.
[0101] "Parameter update" refers to the process of adjusting the recognition model and evaluating the model's internal weights, biases, or other trainable parameters based on the learning data during the learning process.
[0102] "Detection accuracy" refers to the overall performance of the system in terms of accuracy, recall, F-score, and other performance indicators in recognizing bullying-related emotional states and behaviors, as well as determining the likelihood of bullying.
[0103] In this embodiment of the invention, the system mainly consists of a terminal, a server, and an information display device on the user side, which interact with each other through a communication network. The server undertakes the core data processing and model inference tasks, the terminal is responsible for image information acquisition and preprocessing, and the user views the analysis results and inputs feedback information through the display device.
[0104] I. Overall Hardware and Software Composition The terminal can be a smartphone, tablet, desktop computing device, or embedded camera monitoring device. The terminal includes an imaging device (such as a built-in camera or an external camera), a processor, memory, and a network communication module. The terminal runs an operating system (such as a mobile operating system or a general-purpose operating system) and a dedicated acquisition program, which calls the camera driver interface (such as a multimedia-framework-based camera interface) to acquire image information.
[0105] The server can be deployed on a cloud computing platform or in a local data center. It includes a multi-core central processing unit, a graphics processing unit, main memory, and persistent storage. The server runs a server operating system (such as a Linux-based operating system) and deploys image processing libraries (such as OpenCV), video decoding libraries (such as FFmpeg-based libraries), machine learning frameworks (such as general-purpose deep learning frameworks), and a runtime environment for generative artificial intelligence model inference. The server also includes a database management system for storing image features, evaluation values, user feedback, and model parameter version information.
[0106] The user-side information display device can be a smartphone, tablet, or personal computing device running a browser or dedicated application to receive warning messages and response guidelines pushed by the server and display a visual interface.
[0107] II. Implementation Modes of Terminal-Side Program Processing In this invention, the terminal is responsible for converting real-world children's expressions and behaviors into digital image data that can be analyzed by the server. The terminal continuously acquires image information containing the child and their surrounding environment through an imaging device. The terminal uses a video acquisition interface to obtain raw frame data from a camera, which is stored in the terminal's memory in RGB or YUV format.
[0108] The terminal uses an encoding library to encode consecutive frames into a compressed video stream. The terminal can choose to use a video compression standard based on block transform and entropy coding to reduce bandwidth consumption. The terminal automatically adjusts encoding parameters according to network conditions; for example, it uses a higher bit rate and resolution when network conditions are good, and lowers the frame rate and resolution when network conditions are poor, thereby reducing communication load while ensuring the image quality required for recognition.
[0109] During the encoding process, the terminal appends time and identification information to each frame or group of image data. The time information can be a high-precision timestamp, and the identification information can include camera device identification, installation location number, class number, and child's anonymous ID. The terminal encapsulates this information into a metadata structure so that the server can index and aggregate it by time and object during subsequent processing.
[0110] The terminal establishes a connection with the server via secure communication. It utilizes transport layer security protocols for handshake authentication and key negotiation, and employs symmetric encryption algorithms to encrypt the encoded video stream and metadata. During transmission, the terminal performs sequence number marking and retransmission control on data packets to minimize the impact of packet loss on timing analysis. The terminal can also perform local caching, temporarily storing image information for a certain period when the network is interrupted, and retransmitting it after the connection is restored.
[0111] Through the above processing, the terminal converts the continuous visual scene of the real world into a structured, time-aligned video data stream with attached identification information, and sends it to the server in encrypted form, thereby providing high-quality input for high-dimensional temporal analysis on the server side.
[0112] III. Implementation Forms of Server-Side Image and Behavior Analysis Module After receiving image information from the terminal, the server first performs decryption and integrity verification at the network interface layer. Then, it calls the video decoding library to decode the bitstream, restoring it to chronologically ordered frame data. The server maintains a buffer in memory for each video stream, storing the most recently pre-defined frame data along with its time and identification information.
[0113] The server uses an image processing library to perform preprocessing operations on each frame, including resizing, noise filtering, and brightness normalization. The server then uses a face detection algorithm to locate facial regions in each frame. This algorithm employs a convolutional neural network-based detector to perform multi-scale sliding window scanning on the input frames, outputting a set of candidate face boxes containing location coordinates and confidence scores. The server filters valid face regions based on a confidence threshold and then crops and normalizes their size.
[0114] The server extracts facial feature vectors for each face region. The server can use a pre-trained convolutional neural network model as a feature extractor, which contains multiple convolutional layers, pooling layers, and fully connected layers. The server normalizes the cropped facial image and inputs it into this network, reading the feature representations from the intermediate layers as input features for emotion recognition. Subsequently, the server inputs these features into an emotion classification network, which can employ a structure combining convolutional and fully connected layers. The final layer outputs a multi-dimensional probability distribution, corresponding to emotion categories such as "fear," "anger," "sadness," "happiness," and "neutrality." The server determines the current emotional state based on the category with the highest probability, while retaining the probabilities of each category as refined features for subsequent evaluation.
[0115] The server employs a pose estimation algorithm for body regions. It can call a deep neural network-based human keypoint detection model, which accepts whole-frame or local area images as input and outputs the two-dimensional coordinates of multiple body keypoints (such as head, shoulders, elbows, knees, etc.) and visibility scores. The server then arranges the keypoint coordinates of each person across multiple consecutive frames in chronological order, forming a pose time series for that person.
[0116] The server performs action classification on a time series of postures. The server can model this using a temporal convolutional network or a long short-term memory network: the input layer receives a sequence of keypoint coordinates within a given time window, the intermediate layers capture dynamic patterns over time through convolutions or recurrent structures, and the output layer provides action category probabilities, such as "normal communication," "walking," "pushing," "beating," and "surrounding." The server determines the dominant action type for the current time period based on the category with the highest output probability and records the time period and person identifier associated with aggressive or attacked behavior.
[0117] IV. Implementation Forms of Server-Side Feature Fusion and Bullying Assessment Module The server collects time-series data from the emotion recognition and action classification modules to form multimodal features. Within a preset time window (e.g., a few minutes), the server calculates the emotional statistical characteristics of each child, such as the proportion of negative emotion categories, the frequency of emotion transitions, and the mean and variance of emotion intensity. The server also statistically analyzes the frequency and duration of aggressive and passive behaviors, as well as the role distribution of participants when the behaviors occur.
[0118] The server combines the aforementioned statistical features with contextual and object information to form a feature vector. Each dimension of the feature vector corresponds to a specific metric, such as the average probability of feeling "fear" within a certain time window, the proportion of time during which "fear" occurs consecutively beyond a threshold, the number of times a specific person is identified as a victim, and the frequency of occurrence with the same attacker. The server inputs these high-dimensional features into the evaluation model.
[0119] The evaluation model can employ a multilayer perceptron architecture, containing several fully connected layers and nonlinear activation functions. During training, the server uses labeled data (e.g., samples marked "risk of bullying" or "no risk of bullying") to optimize model parameters by minimizing the binary cross-entropy loss function. The server updates weights and biases using a gradient descent-based optimization algorithm, iteratively training until the validation set performance converges. During inference, the server outputs a bullying risk assessment value for each sample, ranging from 0 to 1, with values closer to 1 indicating a higher probability of bullying.
[0120] The server determines the evaluation value based on predetermined criteria. For example, the server sets multiple thresholds to classify risks into "normal," "caution," and "high-risk" levels. The server can dynamically adjust the thresholds based on historical false positive and false negative statistics to achieve a better balance between recall and precision. This determination result, along with the evaluation value, is stored in a database to support subsequent analysis and learning.
[0121] Through the aforementioned feature fusion and evaluation modeling, the server implements a computational method that integrates emotion and action data over time. This method differs from subjective judgment based on human observation and simple rule screening. Instead, it automatically learns complex pattern associations within the computer using high-dimensional features and neural network models, thereby improving the accuracy and robustness of bullying detection.
[0122] V. Implementation Forms of Server-Side Generative Artificial Intelligence Models and Prompt Statement Construction When the server identifies a situation where the likelihood of bullying reaches a predetermined level, it generates structured information describing the current scenario, including the time interval, location, participant roles, main emotional trends, main action types, and risk level. Based on this information, the server constructs prompts for a generative artificial intelligence model.
[0123] The prompts generated by the server are in natural language text. The server can include the following structure: a scenario description, a problem description, and output requirements. For example, the server can construct the following prompt: "Based on surveillance video analysis, the system detected that a student in Class 2, Grade 3, repeatedly displayed expressions of fear between 10:15 and 10:20, and was also subjected to aggressive behavior such as being pushed and shoved by classmates. The bullying risk score is 0.92. Now, we need to generate a response suggestion for the homeroom teacher. As a psychological counseling and education expert, please provide a detailed explanation of how the teacher should communicate with the victimized student after class, how to communicate with the suspected perpetrator, and how to notify the parents. Please provide specific dialogue examples and operational steps." The server can also generate different prompts based on different scenarios, for example: "If a child frequently displays fear or anxiety in the classroom but does not proactively seek help from teachers or parents, as a psychological counselor, what observation and intervention measures would you recommend the teacher to take? Please explain step by step and provide practical suggestions." These prompts are automatically generated by the server based on features and judgment results according to predefined templates and rules. The server will select different tones and levels of detail based on the recipient's attributes (such as homeroom teacher, parent, or school counselor).
[0124] The server sends the pre-constructed prompts to a generative AI model deployed on the same server or a remote inference service. This model can be a large-scale language model based on a self-attention structure, pre-trained on a large-scale text corpus, and fine-tuned on data from the fields of education and psychological intervention. The server takes the prompts as input to the model, which performs forward inference and outputs a natural language text containing coping suggestions, communication techniques, and precautions.
[0125] After receiving the model output, the server performs post-processing on the text based on the evaluation values and recipient attributes. For example, for high-risk scenarios, the server can retain more specific suggestions and emergency response steps; for low-risk scenarios, the server can downplay emergency wording and emphasize observation and communication more. The server can also control the length of the generated content and filter sensitive words to ensure that the output text meets the requirements of the usage scenario.
[0126] By transforming structured analysis results into prompts and using generative artificial intelligence models to generate personalized, context-sensitive natural language suggestions, the server introduces a new text generation process within the computer. This process is not a simple template replacement, but rather uses the expressive power of deep language models to perform language mapping and strategy reconstruction for complex scenarios, thereby improving the matching degree between information presentation and user needs.
[0127] VI. Implementation Forms of User-Side Display and Feedback Users access the server-provided visual interface through an information display device. The server-generated interface includes various views, such as an event timeline displayed by time, a risk profile summarized by object, a sentiment trend chart, a behavior statistics chart, and a text report. On the backend, the server organizes the judgment results, assessment values, sentiment time series, and action classification results into a chart data structure, and the frontend interface renders this data as graphical elements.
[0128] Users can view detailed information for each alert event on the interface, including corresponding video clip screenshots, the emotion probability curve for that period, the type of action detected, and response guidelines text generated by the generative artificial intelligence model. Users can select alert events on the interface and label them with tags such as "confirmed bullying," "no bullying," and "situation unclear," and can also enter text descriptions of the actual investigation and handling results.
[0129] Users provide labeled information to the server through feedback actions on the interface. The server associates this feedback with the current feature vector and evaluation value, and stores it in the learning dataset. The user feedback data provides supervision signals for subsequent model updates, thus forming a complete data loop.
[0130] VII. Implementation Forms of Server-Side Learning and Model Updates The server periodically reads new training data samples from the database. Each sample contains a feature vector, a raw evaluation value, user feedback labels, and relevant contextual information. The server preprocesses the data, including outlier detection, class balancing, and data splitting. The server can employ data augmentation techniques, such as randomly pruning time windows or adding small noise to features, to improve the model's generalization ability.
[0131] During training, the server uses the training set to calculate the loss function. For example, cross-entropy loss is used for the evaluation model, and multi-class cross-entropy loss is used for the emotion recognition and action classification models, respectively. The server uses the backpropagation algorithm to calculate the gradient of the loss function with respect to the parameters of each layer and uses optimization algorithms (such as adaptive learning rate optimization algorithms) to update the parameters. During training, the server monitors the accuracy, recall, and overall metrics on the validation set to prevent overfitting.
[0132] After completing a training cycle, the server evaluates the new model on historical data and a new test set. If the new model outperforms the old model overall, the server replaces the old version in the online service. The server retains the parameters and performance records of the old model so that it can be rolled back in case of anomalies. The server can also adjust the thresholds and weights of the evaluation values based on the frequency of user feedback, such as increasing the weight of certain feature dimensions to more sensitively respond to specific types of bullying patterns.
[0133] Through the above automated learning and model update mechanisms, the server continuously optimizes its internal parameters using new data, thereby continuously improving the system's adaptability to complex scenarios and its detection accuracy. This feedback-driven model iteration differs from manual rule adjustment; instead, it automatically seeks better feature combinations and decision boundaries within the computer through learning algorithms, thus substantially improving computational efficiency and recognition performance.
[0134] VIII. Technical Effects and Causal Relationships The server achieves several technological improvements by employing multi-level feature extraction, temporal modeling, and generative text generation throughout the end-to-end processing chain. By integrating video encoding / decoding optimization and frame sampling strategies on the server side, the number of frames requiring in-depth analysis is reduced, lowering computational and storage pressure while maintaining analytical accuracy. By fusing emotion and action features over time, the server can capture long-term behavioral patterns, not just fleeting states in a single frame, thereby improving the accuracy and stability of bullying risk identification. By constructing prompts and invoking generative AI models to generate context-sensitive response text, the server transforms complex analysis results into easily understood and executable natural language suggestions, enabling users to quickly take effective measures while reducing the time spent manually drafting solutions.
[0135] By introducing a user-feedback-driven learning mechanism, the server uses false positives and false negatives to update the model, forming an adaptive feedback loop. This mechanism enables the system to automatically adjust its internal parameters as data distribution and scenarios change, eliminating the need for frequent manual parameter tuning and thus achieving self-optimization capabilities in the computing system. These improvements go beyond simply automating human workflows; they involve using specific neural network structures, feature construction methods, loss function design, and threshold adjustment strategies to create an algorithmic process within the computer that can be directly executed by non-human intuitive thinking. Therefore, this represents an improvement to computer technology itself.
[0136] IX. Alternative Solutions and Modified Implementation Forms In practical implementation, servers can use different combinations of network structures and algorithms. For example, emotion recognition models can employ multi-branch convolutional networks or structures with attention mechanisms to highlight facial regions relevant to emotion judgment; action classification models can use graph convolutional networks to represent key points of the human body and their connections as a graph structure to better model interactions between joints. Evaluation models can employ ensemble learning methods based on gradient boosting trees, rather than being limited to neural networks.
[0137] During image acquisition, the terminal can add a local pre-screening function. For example, the terminal can use a lightweight model to perform preliminary analysis on frames, sending only frames suspected of containing abnormal expressions or movements to the server, thereby significantly reducing network load at the source. When receiving data, the server can dynamically adjust the frame sampling rate based on the channel load, achieving adaptive resource scheduling.
[0138] User feedback can be input in various ways, such as voice-to-text or quick annotation using template options. The server can perform confidence modeling on user feedback, assigning different weights to different user groups to improve the quality of the learning data.
[0139] Through the above-described various implementation methods and alternatives, the present invention provides a scalable and easily integrated system architecture that enables terminals, servers, and users to collaboratively achieve a complete technical process from image acquisition, temporal analysis, risk assessment, natural language generation to feedback learning, thereby enabling efficient and accurate technical monitoring and intervention support for child bullying risks in the real world.
[0140] use Figure 11 The processing flow is explained.
[0141] Step 1: The terminal uses an imaging device to collect image information.
[0142] The terminal's input is a continuous stream of raw video signals from a camera. The terminal uses the camera driver interface provided by the operating system to acquire raw image frames containing the child's expressions and behaviors from the imaging device at a preset resolution and frame rate. The terminal adds time and recognition information to each frame, writing timestamps, device identifiers, location identifiers, etc., into the corresponding metadata structure. Through this data processing, the terminal converts the unstructured continuous optical signal into digital image frames ordered by time and containing identifier metadata. The output is a sequence of raw image frames with time and recognition information.
[0143] Step 2: The terminal compresses, encodes, and packages the image frames.
[0144] The terminal's input is the raw image frame sequence and its metadata output from step 1. The terminal calls a video coding library to perform compression encoding on the continuous image frames, converting the high-volume raw data into a video stream. Based on the current network conditions, the terminal selects bitrate, resolution, and frame rate parameters, and performs block partitioning, transformation, and entropy encoding on the image pixels to reduce the data volume. Simultaneously, the terminal encapsulates time information and recognition information into container-format headers or side information to align the data with the metadata. The output is a sequence of data packets containing the video stream and corresponding metadata.
[0145] Step 3: The terminal sends data to the server via secure communication.
[0146] The terminal's input is the data packet sequence generated in step 2. The terminal first establishes an encrypted communication connection with the server, verifies the server certificate, and generates a session key. The terminal encrypts the data packets using a symmetric encryption algorithm and adds a sequence number and checksum to each packet to support out-of-order reordering and integrity verification. The terminal uses its network communication module to send the encrypted data packets one by one to the server address via a secure protocol, and retransmits them in case of packet loss or timeout. Through this data processing, the terminal reliably transmits local video data to the remote server. The output is an encrypted data stream transmitted over the network to the server.
[0147] Step 4: The server receives and decrypts the video data stream.
[0148] The server receives encrypted data streams from the terminal. At the network interface, the server receives data packets, performs decryption operations using the session key, and recovers the original video stream and metadata. The server performs integrity checks and sequence reassembly on each data packet, restoring the correct time sequence based on the sequence number. Through these data operations, the server converts the scattered encrypted data packets into a coherent time-ordered stream and extracts time and identification information. The output is a decodeable video stream along with the corresponding time and identification metadata.
[0149] Step 5: The server decodes the video stream and performs frame preprocessing.
[0150] The server's input is the video stream and metadata obtained in step 4. The server calls the video decoding library to perform decoding operations on the stream, restoring it into continuous image frame data. The server samples the frame sequence according to a preset strategy, such as selecting a certain number of frames per second, to balance computational load and temporal resolution. The server performs size scaling, color space conversion, and noise filtering on the sampled frames, normalizing the pixel matrix to the numerical range required by the model. Through these data processing steps, the server obtains standardized image frames that retain key visual information while reducing redundancy. The output is the preprocessed image frame sequence along with its time and recognition information.
[0151] Step 6: The server detects facial regions and extracts facial features in the frame.
[0152] The server's input is the preprocessed image frame output from step 5. The server calls a face detection algorithm from an image processing library to search for facial locations in each frame, obtaining the coordinates and confidence scores of several facial regions. The server filters low-confidence candidate boxes based on a threshold and crops and uniformly scales the remaining regions. Subsequently, the server inputs the cropped facial image into a pre-trained convolutional feature extraction network, performs forward propagation, and calculates multiple convolutional and pooling operations to obtain high-dimensional feature vectors. Through this series of data operations, the server converts the two-dimensional pixel matrix into a vector representation of expression patterns. The output is the feature vector corresponding to each facial region, along with its time and recognition information.
[0153] Step 7: The server estimates the emotional state based on facial features.
[0154] The server's input is the facial feature vector output from step 6. This feature vector is fed into an emotion classification model containing several fully connected layers and non-linear activation units. The server calculates scores for each emotion category using matrix multiplication and activation functions, and then converts these scores into probability distributions using a normalization function. The server determines the current emotion category based on the maximum probability, while retaining the complete probability vector as fine-grained information. Through these data processing steps, the server maps abstract features to specific emotional states. The output is the emotion category and probability distribution for each face at each time point.
[0155] Step 8: The server detects body regions in the frame and performs pose estimation.
[0156] The server's input is the preprocessed image frame output from step 5. The server calls a pose estimation model to analyze the entire frame or region of interest, outputting the coordinates and confidence scores of keypoints for multiple people's bodies through convolution operations and keypoint regression. The server filters valid keypoints based on the confidence scores and assigns a unique identifier to each identified person, constructing a mapping table between people and keypoints. The server concatenates the keypoint coordinates of the same person in consecutive frames in chronological order, forming a pose time series. Through this data processing, the server encodes the body morphology in the original image into computable temporal coordinate data. The output is the keypoint time series and its identifier information for each person.
[0157] Step 9: The server performs action classification on the posture time series.
[0158] The server's input is the keypoint time series output from step 8. The server selects a fixed-length time window and uses the keypoint coordinate sequence within that window as the model input. The server invokes a temporal action classification network, calculating temporal features through multiple layers of temporal convolutions or recurrent units to represent the sequence in high dimension. The server calculates the probability of each action category at the output layer, such as "normal communication," "walking," "pushing," "beating," and "surrounding." The server determines the dominant action for the current window based on the highest probability and records the action label and probability along with the corresponding person's identifier. Through these data processing steps, the server converts continuous pose change patterns into discrete behavior categories. The output is the action category and probability for each time window and person.
[0159] Step 10: The server constructs time-series features of emotions and behaviors.
[0160] The server takes as input the emotional state sequence obtained in step 7 and the action category sequence obtained in step 9, aligning the two types of information according to time and person identification. Within a predetermined time window, the server statistically analyzes the distribution of emotional states, the duration of continuous negative emotions, and the frequency of emotion transitions; it also statistically analyzes the frequency and average duration of aggressive and passive behaviors, as well as the relationship patterns between participants. The server combines these statistical values with contextual and object information to form a high-dimensional feature vector. Through this feature construction operation, the server integrates the original time-series classification results into a numerical representation suitable for use in the evaluation model. The output is the set of feature vectors corresponding to each time window.
[0161] Step 11: The server assesses the likelihood of bullying based on the feature vector.
[0162] The server's input is the feature vector output from step 10. The server feeds these features into an evaluation model, typically a multi-layer neural network or other supervised learning model. The server computes an intermediate representation using linear transformations and non-linear activations, obtaining an evaluation value in the range of 0 to 1 at the output layer, representing the likelihood of bullying. The server compares the evaluation value with predetermined criteria, such as one or more thresholds, to determine whether the current time window belongs to a category such as "normal," "attention," or "high risk." Through this series of data operations, the server transforms multi-dimensional statistical features into a single risk measure and judgment result. The output is the evaluation value and bullying likelihood judgment result for each time window.
[0163] Step 12: The server generates a structured scene description based on the judgment result.
[0164] The server's input consists of the evaluation values and judgment results output from step 11, along with the context information of the corresponding time window. The server filters out time periods where the evaluation values reach a predetermined level, combining the time range, location, participant roles, main emotional trends, main action types, and risk levels into a structured scene description object. The server organizes and encodes these fields to form a highly readable intermediate representation. Through this data processing, the server converts the internal numerical judgment results into a high-level semantic description for natural language generation. The output is a scene description object used to generate prompt statements.
[0165] Step 13: The server automatically generates prompts based on the scenario description.
[0166] The server's input is the scene description object output in step 12. Based on preset templates and rules, the server fills in time, location, emotion, action, and evaluation value into a natural language sentence, generating a prompt statement for the generative AI model. The server can adjust the wording and level of detail based on the recipient's attributes, for example, using different titles and information depths for teachers or parents. Through this text construction operation, the server maps structured data into text input that can be understood by the language model. The output is a prompt statement containing a scene description and a problem description.
[0167] Step 14: The server invokes a generative artificial intelligence model to generate response text.
[0168] The server's input is the prompt statement generated in step 13, which is passed as an input sequence to the deployed generative AI model. The model maps words to vectors through an embedding layer, computes context-sensitive representations through a multi-layer self-attention network, and generates natural language text word-by-word in the output layer according to a probability distribution. During inference, the server controls the output length and decoding strategy to avoid repetition or deviation from the topic. Through this sequence generation operation, the server obtains a natural language text containing the warning content and corresponding response guidelines. The output is a warning message and response guidelines specific to the current scenario.
[0169] Step 15: The server performs post-processing and personalization adjustments on the generated text.
[0170] The server's input consists of the natural language text output from step 14, the evaluation value from step 11, and the recipient attributes. Based on the risk level, the server decides whether to enhance the emergency warning statement and selects appropriate suggested paragraphs based on the user's role. The server can also perform sensitive word filtering, sentence length optimization, and formatting to ensure the text is semantically rigorous and conforms to usage guidelines. Through this post-processing, the server adjusts the general generated result into personalized content more tailored to the specific user and scenario. The output is the final warning message and response strategy text.
[0171] Step 16: The server pushes warning information to the terminal or user display device.
[0172] The server's input consists of the warning message and response strategy text output from step 15, along with associated assessment values and analysis summaries. The server invokes a push notification service or interface to package this content into a notification message or report. The server then sends the message to the corresponding user's terminal or browser frontend, allowing the user to view detailed information on the interface. Through this data transmission and organization, the server presents the calculation results to the actual user. The output is the warning notification and detailed report data that can be displayed on the user's device.
[0173] Step 17: Users can view the alert details and provide feedback.
[0174] The user's input is the warning notification and report presented on the terminal or display device in step 16. The user reads the assessment values, emotional trends, action categories, and response guidelines text through a graphical interface. Based on the actual investigation results, the user selects tags on the interface, such as "Bullying Confirmed," "False Alarm," or "Continued Observation Required," and can also enter text descriptions. Through this interactive operation, the user transforms subjective judgments and on-site information into structured feedback data. The output is a feedback record with tags and descriptions.
[0175] Step 18: The server receives feedback and updates the learning dataset.
[0176] The server's input consists of the feedback records output from step 17, along with the feature vectors, evaluation values, and decision results for the corresponding time windows. The server associates and stores this data in the learning dataset. It then performs quality checks and normalization on the newly added data to prepare for subsequent model training. Through this data aggregation operation, the server integrates user feedback into a sample set that can be used for supervised learning. The output is an updated learning dataset containing the new samples.
[0177] Step 19: The server performs model retraining and parameter updates based on the learning dataset.
[0178] The server's input is the learning dataset output from step 18. The server divides the samples into training and validation sets, loading the emotion recognition, action classification, and evaluation models into trainable states. During the training loop, the server calculates the loss function, compares the model output with user feedback labels, calculates the gradient using backpropagation, and updates the model weights using an optimization algorithm. During training, the server adjusts hyperparameters based on validation set performance, and after training, deploys the better-performing model to the online inference module. Through this model update operation, the server gradually reduces errors and improves the consistency between the evaluation values and the actual situation. The output is the updated model parameters and the new decision policy.
[0179] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0180] In existing school and home guardianship scenarios, the detection of the risk of bullying against children mostly relies on simple rule matching, threshold alarms, or static emotion recognition models. These technical solutions typically only assess limited features within a single frame or short time period, lacking the ability to comprehensively analyze long-term behavioral patterns, complex scene semantics, and the relationships between multi-dimensional features, thus leading to the following technical problems: (1) When processing multi-source visual data of children’s facial expressions and movements, computers lack effective data structuring and time window statistical mechanisms, and cannot stably and accurately extract key features reflecting bullying situations from continuous images, making the detection results susceptible to noise and short-term anomalies, resulting in insufficient robustness. (2) Traditional identification methods based on fixed model parameters often give a "risk or no risk" judgment directly through preset algorithm logic. The calculation process lacks interpretability, and the computer system cannot adaptively adjust the judgment criteria according to the complex environment. It is difficult to make comprehensive inferences on high-dimensional features such as emotion ratio, defensive posture, and contact behavior of others, which limits the practicality of the model in real-world scenarios. (3) Generative AI models have powerful reasoning and text generation capabilities, but in the existing technology, the server side usually does not build an automatic mapping mechanism between structured features and prompt statements for the specific task of bullying detection. As a result, generative AI models are only used as general question answering tools. Computer systems cannot efficiently and stably call generative AI models for pattern reasoning and risk assessment in the end-to-end data processing pipeline. (4) In real-time or near-real-time monitoring scenarios, the server lacks an integrated processing flow control mechanism from image acquisition, feature extraction, generative artificial intelligence model invocation to alarm generation, resulting in a large overall computing link latency, which cannot meet the needs of child safety monitoring that requires rapid response. (5) In the existing system, there is insufficient utilization of historical data such as bullying detection results, alarm levels and guardian response information. The server side usually does not build a closed-loop feedback update mechanism for the generation conditions of the model and prompt statements, making it difficult for the computer system to continuously optimize the recognition algorithm and reasoning strategy based on the actual usage process, thus limiting the improvement of detection accuracy under long-term operation of the system.
[0181] Therefore, a new computer implementation is needed to improve the computer's data representation, reasoning, and online adaptability in child bullying risk detection tasks by using structured processing of continuous image data on the server side, automatic construction of prompts based on generative artificial intelligence models, and a feedback-driven dynamic update mechanism. This would enhance the system's detection accuracy, real-time performance, and interpretability in complex real-world scenarios.
[0182] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.
[0183] In this invention, the server includes: a device for receiving image information and corresponding attribute information from a terminal that collects children's image information; a device for converting the image information into time-series image information and performing target recognition processing, facial expression recognition processing, and action recognition processing on the image information to extract feature quantities related to the child's emotional state and body movements; a device for summarizing the feature quantities according to a predetermined time unit to generate structured data as statistical information, the statistical information including at least the proportion of emotional states occurring within the time unit, the duration of defensive postures, and the number of contact behaviors from others; a device for automatically generating prompt statements for inputting into a generative artificial intelligence model based on the structured data and text information describing the structured data, and inputting the prompt statements and the text information into the generative artificial intelligence model; and a device for... The device outputs a reasoning result to assess the likelihood of the child being bullied, the reasoning result including at least a bullying risk indication expressed in numerical form and a reasoning explanation for the bullying risk indication, the assessment including a device for determining the bullying risk indication as a warning level; a device for sending a warning message containing the assessment result and the reasoning explanation to the terminal when the warning level meets a predetermined benchmark, and causing a warning message for notification to the guardian or education-related party based on the warning message; and a device for generating countermeasure information including coping strategies and intervention behavior candidates for the child based on the assessment result, and optionally, storing the reasoning result, the assessment result, the warning level, and the response information corresponding to the warning message as historical information, and updating the target recognition processing, the facial expression recognition processing, the action recognition processing, and the generation conditions of the prompt statement based on the historical information. This allows for the conversion of continuous children's images into a structured representation suitable for generative artificial intelligence model inference at the computer level. By organically connecting visual features with the language model inference process through automatically constructed prompts, the server can comprehensively assess bullying risks in real-time or near real-time and output interpretable warning results. At the same time, it can form a feedback loop by utilizing historical interaction data, thereby significantly improving the computer system's data processing capabilities, inference accuracy, response speed, and adaptive optimization capabilities in child bullying detection tasks.
[0184] "Terminal" refers to an information processing device used to collect children's image information and send the image information and attribute information to the server, including but not limited to electronic devices with camera function and network communication function.
[0185] "Image information" refers to time-series visual data about children and their surroundings that are acquired through imaging devices and can be processed by computers, including but not limited to video data and image frame data extracted from videos.
[0186] "Attribute information" refers to additional data associated with image information and used to assist in analysis, including but not limited to acquisition time, acquisition location, device identification, scene type, and identification information related to children.
[0187] "Time-series image information" refers to a data sequence composed of multiple image frames arranged in chronological order, used to characterize a child's continuous behavior and facial expression changes over a period of time.
[0188] "Target recognition processing" refers to the computational processing performed on image information to detect and identify one or more object regions in an image, including but not limited to the detection and localization of people, body parts and related objects.
[0189] "Facial expression recognition processing" refers to the analysis of a child's facial area to determine the category of emotion and its confidence level, including but not limited to the recognition of emotional states such as happiness, fear, sadness, and anger.
[0190] "Motion recognition processing" refers to the computational processing that analyzes a child's body posture and its movement patterns over time to identify behavioral characteristics such as defensive postures, being pushed, or being pulled.
[0191] "Features" refer to numerical representations obtained from target recognition processing, facial expression recognition processing, and action recognition processing, used to quantitatively describe children's emotional states and physical movements, including but not limited to emotion confidence, key point coordinates, posture labels, and behavior labels.
[0192] "Preset time unit" refers to a time interval that is pre-set in the system as the basis for statistics and analysis, including but not limited to several seconds, several tens of seconds, or several minutes.
[0193] "Statistical information" refers to the summary data calculated based on characteristic quantities within a predetermined time unit to reflect the overall situation of emotional state, action pattern and interaction behavior, including proportion, duration and frequency.
[0194] "Structured data" refers to data formats in which statistical information is organized according to a predetermined data format, making it easy for computers to store, retrieve, and further process it, including but not limited to key-value pairs, tabular formats, or other parsable data structures.
[0195] "The proportion of occurrence of an emotional state" refers to the percentage of time or frequency in which a specific emotional state appears in all image frames or all detection results within a predetermined time unit.
[0196] "Duration of defensive posture" refers to the cumulative length of time within a predetermined time unit during which a child is judged to be in a defensive body posture.
[0197] "Number of contact behaviors from others" refers to the number of times the system detects contact behaviors (including pushing, pulling, etc.) that others have engaged in with the child within a predetermined time unit, based on the action recognition results.
[0198] “Textual information” refers to natural language text used to describe the content and context of structured data, including but not limited to verbal descriptions of statistical results, time windows, and behavioral patterns.
[0199] "Generative AI models" refer to AI models that can generate output content based on input structured data and prompts, including but not limited to deep learning-based language models or multimodal models.
[0200] "Prompt statements" refer to natural language text input into a generative artificial intelligence model that indicates the model's task objectives, constrains the output format, and provides contextual information.
[0201] "Inference results" refer to the analysis results output by the generative artificial intelligence model based on prompts and related input data, including bullying risk indicators and explanatory information related to the risk.
[0202] "Bullying risk indicator" refers to a quantitative indicator that represents the likelihood of a child being bullied in the form of a numerical value or a sortable ranking, and is used by the server to conduct risk assessment and determine the warning level.
[0203] "Explanation of Reasons" refers to the natural language explanation text provided in response to bullying risk indications, which describes the basis for risk assessment and related circumstances.
[0204] "Assessment results" refers to the judgment information obtained by the server after comprehensively processing the bullying risk indications and reasons contained in the reasoning results to determine the likelihood of a child being bullied.
[0205] "Warning level" refers to classification or grading information that indicates the severity of risk based on assessment results, including but not limited to levels such as normal, caution, warning, and severe warning.
[0206] "Warning message" refers to data generated by the server and sent to the terminal when the warning level meets the predetermined benchmark. It includes the assessment results and explanations and is used to trigger a notification to the guardian or education-related party.
[0207] "Warning messages" refer to notifications generated by the terminal based on warning information and presented to guardians or education-related parties, including but not limited to push notifications, text messages, emails, or pop-up windows.
[0208] "Countermeasure information" refers to guidance data generated by the server based on the assessment results, which includes response strategies and intervention candidates for children, and is used to assist guardians or education stakeholders in developing handling measures.
[0209] "Guardian" refers to an individual or organization that has guardianship responsibility for a child, including but not limited to parents, other legal guardians, and related caregivers.
[0210] "Education stakeholders" refers to entities involved in the management and education of children in educational settings, including but not limited to teachers, school administrators, and other staff of educational institutions.
[0211] "Historical information" refers to past data accumulated and stored during system operation that is related to reasoning results, evaluation results, warning levels, warning messages, and the response behavior of guardians or education stakeholders.
[0212] "Generation conditions" refers to the rules, parameters, and constraints set when automatically generating prompts, including but not limited to text templates, feature selection strategies, and output format requirements.
[0213] The embodiments of the present invention embody the information processing system described in Appendices 1 to 3 as a technical solution for collaborative operation between the server, terminal, and user. In the following description, the subjects are limited to "server," "terminal," and "user," and the implementation of the present invention is described in detail, focusing on specific hardware, software, data structures, and algorithm flows.
[0214] I. Overall System Composition A server includes a processor, memory, network interfaces, and optional graphics processing units. Servers can be implemented using computing devices running a general-purpose operating system, such as a rack-mounted computing unit running Linux. The software aspects of a server include: The server includes a network service module for handling network requests, a visual analysis module for performing image processing, a feature modeling module for performing feature statistics and structuring, an inference module for constructing prompt statements and calling generative artificial intelligence models, a database module for storing and retrieving data, and a history management module for log and feedback processing.
[0215] The server can use open-source image processing libraries and deep learning frameworks in the visual analysis module. For example, the server can use OpenCV to perform video decoding, frame extraction and basic image operations, and use TensorFlow or PyTorch to load and execute deep neural network models for face detection, expression recognition and action recognition.
[0216] The terminal includes a camera device, a processor, a memory, and a network communication module. The terminal can be a smartphone, a tablet computer, or a fixed video surveillance device. At the software level, the terminal includes a data acquisition application running on a mobile operating system. This application is responsible for controlling the camera, performing local preprocessing of the acquired video data, and sending the data to the server via a secure communication protocol.
[0217] Users can view the results returned by the server through the terminal. In the graphical user interface, users can view video clips, risk scores and explanatory text, and can confirm or provide feedback on alarms.
[0218] II. Program Modules and Data Processing Methods 1. Server-side software modules and data structures The server processes data sequentially using multiple data structures throughout the processing chain.
[0219] When the server receives video information, it stores the video file's path and attribute information as metadata entries in a relational database or document database. This metadata includes the device identifier, acquisition timestamp, location encoding, video file storage location, and subsequent analysis status identifier.
[0220] When decoding a video, the server divides the video into evenly spaced image frames. The server assigns a frame identifier and a timestamp to each frame and organizes them into a time-series list for subsequent sequential processing.
[0221] When extracting features, the server creates a frame-level feature object for each frame. This object includes: an emotion probability vector for the facial region, an array of coordinates for key points in the human pose, a set of action labels, and auxiliary information corresponding to the frame timestamp. During statistical processing, the server aggregates frame-level features from consecutive seconds or tens of seconds to form time-window-level statistical information. This statistical information is structured as key-value pairs, such as "fear_ratio", "defensive_posture_duration", and "aggressive_contact_count".
[0222] When generating the prompt, the server converts the structured data into a natural language description. The server then uses a rule template module to fill the structured data into a preset sentence pattern, forming a readable Chinese description, which is then combined with the task description to form a complete prompt text.
[0223] When a server invokes a generative AI model, it communicates with the model service deployed locally or in the cloud using HTTP or RPC protocols. The server includes prompt text and necessary context information in the request, and receives bullying risk indications and explanations generated by the model in the response.
[0224] 2. Terminal-side software modules and data processing methods In this invention, the terminal is not only responsible for data acquisition, but also undertakes some preprocessing tasks.
[0225] In the data acquisition application, the terminal controls the camera device through the operating system's camera interface to generate video stream data at a fixed resolution and frame rate. The terminal encodes the video locally, for example, using H.264 or H.265 encoding formats, to reduce storage and transmission bandwidth.
[0226] Before uploading, the terminal can compress and segment the video. The terminal can also attach device identifiers, logical classroom numbers or home location numbers, video start and end times, etc. to the metadata so that the server can accurately restore the scene.
[0227] When receiving analysis results from the server, the terminal parses the JSON or other structured response returned by the server into a local object. The terminal generates graphical elements in the interface presentation layer, including risk level color markings, text descriptions, and replay buttons, and decides whether to trigger the system notification service based on the assessment results, alerting the user in the status bar or pop-up window.
[0228] III. Visual Recognition Model and Algorithm Implementation In this invention, the server incorporates various deep learning structures to achieve detailed analysis of children's facial expressions and movements.
[0229] The server can use a classification model based on convolutional neural networks for facial expression recognition. The model takes a normalized facial image tensor as input and outputs a probability distribution vector for each emotion. The server is designed with a multi-layered structure of convolutional, pooling, and fully connected layers. Cross-entropy loss can be used as the loss function, and gradient descent-based adaptive optimization can be employed for weight updates.
[0230] The server can use models based on convolutional neural networks and keypoint regression for human pose estimation. The model outputs the positions of human keypoints in the image coordinate system. Based on the geometric relationships between keypoints, the server infers the existence of action patterns such as defensive postures and being pushed or pulled through rules or small classification sub-models.
[0231] For action recognition, the server can use temporal convolutional networks or long short-term memory-based models, taking pose information from several consecutive frames as input and outputting action category labels. Based on these labels, the server counts the number of attack actions and their duration within a predetermined time window.
[0232] When training the aforementioned visual model, the server performs offline training using labeled sentiment and behavior datasets. During the training phase, the server defines the loss function as the difference between the labeled categories and the model's predictions. The server iteratively updates the network weights using the backpropagation algorithm and optionally employs data augmentation strategies, such as image flipping, random cropping, and illumination enhancement, to improve the model's robustness to changes in real-world scenes.
[0233] The server fixes the model parameters during the inference phase and only performs forward propagation to ensure inference speed. During inference, the server uses batch processing and GPU acceleration to improve the throughput and real-time performance of frame-level analysis.
[0234] IV. Generative Artificial Intelligence Models and Prompt Statement Construction In this invention, the server uses a generative artificial intelligence model as the upper-layer inference module. This model can be a language model or a multimodal model based on the Transformer architecture.
[0235] The server uses prompts at the input end to convert visual statistical information and scene descriptions into natural language text, which is then combined with task instructions to form a complete input. The server's prompts may include system role settings, task instructions, and data descriptions.
[0236] In this implementation, the server can use the following example of a prompt statement: Example of a prompt statement 1: "You are a generative AI model specifically designed to analyze children's emotions and behaviors. Based on the statistical characteristics and scenario description below, please determine whether the child is likely to be bullied, and provide a risk score from 0 to 1 along with a brief reason. Scenario description: In the last 30 seconds, the child displayed a fearful expression 70% of the time; the child maintained a defensive posture for 20 seconds; the system detected 6 suspected aggressive contact behaviors. Please output the results in JSON format, including the fields bullying_risk (numerical value) and explanation (textual description)." Example of a prompt statement 2: "Based on the characteristics analyzed in the video below, determine whether the child is likely to be bullied, and explain the main basis to the teacher in concise language. The characteristics include: the percentage of fearful expressions, the duration of defensive actions, and the frequency of aggressive contact." The server employs a rule-based template generation strategy in its prompt statement construction module: based on the value of each field in the structured data, the server fills the corresponding language slots to generate sentences in a uniform format. This mapping from structured data to text avoids subjective human intervention, ensuring that the input received by the generative AI model is stable and standardized, which is beneficial for improving the consistency and interpretability of inference results.
[0237] After invoking the generative artificial intelligence model, the server receives the text output by the model and parses the numerical risk score and explanation from it. According to Note 1, the server maps the risk score to a warning level.
[0238] During model training or fine-tuning, the server can employ supervised learning methods. It utilizes historical labeled cases to construct a mapping between input prompts and expected outputs, using cross-entropy loss or sequence-to-sequence loss functions, and employing gradient descent-based optimization algorithms to update model parameters. During training, the server can use techniques such as teacher coercion and label smoothing to stabilize the training process and improve generalization ability.
[0239] V. Technical Effects and Improvements in Computer Technology The server uses time-window aggregation and structured data modeling to compress the original high-dimensional, noisy video frame data into a small number of key statistical features. This data structuring approach allows generative AI models to process only concise text descriptions, thus significantly reducing computational overhead and communication burden during inference.
[0240] The server employs a layered processing architecture, decoupling the underlying visual recognition computations from the upper-level language inference. This allows the visual model to be optimized independently, while the generative AI model focuses on pattern interpretation and risk assessment. This layered structure enables the system to expand to new scenarios by updating only certain modules, improving the system's maintainability and scalability.
[0241] The server standardizes the input text using prompt templates, ensuring that the generative AI model receives consistent and structurally similar inputs across different time windows, terminals, and environments. This reduces the variance of the model's output and improves overall evaluation accuracy.
[0242] The server records inference results, evaluation results, and user feedback through a historical information management module. Within a certain period, the server adjusts the confidence threshold, time window length, and descriptive granularity of the visual model based on this historical data, thus forming a feedback loop. This mechanism allows the system to automatically correct false alarms and missed alarms during continuous operation, optimize the overall computation process and parameter selection, and demonstrates improved adaptive capabilities of the computer system.
[0243] The server employs GPU-accelerated visual model inference and a batch processing strategy to process video clips uploaded from multiple terminals in parallel. Combined with lightweight representation of structured data, the system can maintain low latency and high throughput even in high-concurrency scenarios, achieving a significant improvement in processing speed.
[0244] By incorporating multiple dimensions of features (emotional proportion, duration of defensive posture, number of attack contacts, etc.) into the prompt statements, the server enables the generative AI model to consider multiple statistical indicators simultaneously during inference, rather than relying on a single threshold judgment. This multi-dimensional comprehensive reasoning approach surpasses traditional judgment methods based on a single rule, effectively reducing the false positive rate and improving the ability to identify complex bullying patterns.
[0245] VI. Differences from traditional human work and rule systems In this invention, the server does not simply program the manual observation process. Instead, it uses a deep neural network to represent visual signals in a high dimension, encodes temporal information through time window statistics, and performs semantic-level reasoning on the structured description through a generative artificial intelligence model. This enables large-scale, multi-terminal real-time analysis that is difficult for humans to perform manually.
[0246] The server does not rely on fixed manual rules in action recognition. Instead, it automatically discovers the implicit correlation between key point movement patterns and bullying behavior through supervised learning on a large amount of labeled data. When generating prompts, the server adopts an asymmetric information compression strategy, retaining only the statistical features that contribute most to the judgment. These features are not the simple counts commonly used in traditional monitoring systems, but high semantic features based on the output of deep models.
[0247] The server employs a unified data flow design and modular structure throughout the entire processing chain, ensuring that the input and output of each module have clearly defined data structures. This avoids the inefficient methods of relying on large amounts of unstructured logs and manual interpretation in traditional systems, thereby fundamentally improving data management and computational efficiency.
[0248] VII. Alternative Solutions and Modified Implementation Methods The server's visual model can be replaced with a neural network based on other architectures, as long as it can generate frame-level emotion and action features. For example, the server can use a visual Transformer-based model instead of a convolutional network to improve its ability to model complex scenes.
[0249] In the generative artificial intelligence model part, the server can use language models or multimodal models of different sizes. When the server is resource-constrained, it can use models with smaller parameter scales to reduce inference costs. In scenarios with high accuracy requirements, the server can use larger-scale models to obtain more detailed inference capabilities.
[0250] The server can use a learnable text generation module to generate prompts instead of fixed rule templates, so that the prompts can automatically adjust the wording and expression based on historical effects while maintaining structural constraints.
[0251] In some embodiments, the terminal can execute part of the facial expression or action recognition model locally, sending only the extracted features to the server, thereby significantly reducing the amount of video data uploaded, alleviating network load, and still maintaining basic risk assessment functions when network conditions are poor.
[0252] Users can adjust system parameters according to their needs in different application scenarios. For example, in a kindergarten scenario, the time window can be shortened to improve response speed, while in a large campus scenario, the time window can be lengthened to reduce false alarms. The server achieves flexible adaptation through the same modular structure under different scenario configurations.
[0253] By combining the aforementioned hardware configuration, software modules, data structure design, and deep learning and generative artificial intelligence models, the system of this invention achieves an end-to-end technical link from raw video to interpretable risk assessment results in the task of detecting child bullying risks. This link demonstrates substantial improvements in computer technology itself in terms of improving detection accuracy, reducing false alarm rates, enhancing real-time performance, and improving data management.
[0254] use Figure 12 The processing flow is explained.
[0255] Step 1: The terminal receives user settings and initializes the collection parameters.
[0256] The terminal takes user-inputted or preset parameters from the application interface as input, including video resolution, frame rate, acquisition time period, upload interval, and network access method. Based on this input, the terminal writes a configuration file or memory configuration structure to local storage, serving as the control basis for subsequent acquisition and upload modules. The terminal calls the operating system's camera interface to initialize the camera hardware, including turning on the camera device, setting the resolution and frame rate, and detecting the current network status. The terminal outputs the status information of the initialized camera and network modules, as well as acquisition configuration data available for subsequent steps.
[0257] Step 2: The terminal collects children's image information and performs local encoding and segmentation.
[0258] The terminal continuously receives raw image frame data from the camera hardware, using the initialized camera module and acquisition configuration data as input. The terminal encodes the continuous image frames, for example, by calling a hardware encoder or software encoding library to compress the raw frames into an H.264 or H.265 video stream. The terminal divides the video stream into multiple segment files according to a preset time length (e.g., 30 seconds) and generates a filename and start / end timestamps for each segment. The terminal can optionally adaptively adjust the video bitrate to match the current network bandwidth. The terminal outputs a series of locally cached video segment files and their associated time information.
[0259] Step 3: The terminal generates attribute information and packages it for uploading.
[0260] The terminal takes video clip files and acquisition configuration as input, and generates attribute information corresponding to each video clip, including device identifier, acquisition start time, acquisition end time, location identifier, and possible child anonymity identifier. The terminal organizes this attribute information into a structured data object and packages the video clip files and attribute information into an upload request data packet. The terminal calls the network communication module to establish a connection to the server via HTTPS or other secure protocols. The terminal's output is an upload request sent to the server, containing image information and attribute information.
[0261] Step 4: The server receives upload requests and stores the original image information.
[0262] The server takes the upload request sent by the terminal as input and reads the binary data and corresponding attribute information of the video segment from the network interface. The server saves the video segment in the file system or object storage, generates a storage path, and creates a metadata record in the database, including the device identifier, timestamp, location, video path, and processing status flag. After a successful write, the server updates the processing queue, adding the video task to the list to be analyzed. The server outputs a video task record registered in the database and the raw video data stored on persistent media.
[0263] Step 5: The server decodes the video and extracts time-series image frames.
[0264] The server takes the stored raw video data path and task records as input, and calls a video processing library (such as OpenCV or an FFmpeg-based library) to open and decode the video file. The server uniformly extracts image frames from the video stream at a preset frame rate (e.g., 5 frames per second), converting each frame into an image matrix of uniform size and color space. The server assigns a frame number and an accurate timestamp to each frame and stores this frame-level data in a memory buffer or temporary storage, forming a time-ordered frame queue. The server's output is the corresponding video frame sequence and its metadata information for a given period of time.
[0265] Step 6: The server performs face detection and cropping of the facial area.
[0266] The server takes time-series image frames as input, calls a face detection model or algorithm, and performs object detection on each frame. The server calculates the bounding box coordinates of the detected face regions and crops a facial sub-image from the original image based on these coordinates. The server normalizes the size and pixel values of the facial sub-image to fit the input format of the expression recognition model. The server records the association between each facial sub-image and its corresponding frame number and timestamp. The server's output is a set of preprocessed facial image tensors and their index information.
[0267] Step 7: The server performs facial expression recognition processing and generates emotional feature data.
[0268] The server takes preprocessed facial images as input and feeds them into an expression recognition neural network deployed on a deep learning framework. Using forward propagation through convolutional and fully connected layers, the server outputs a probability vector for each facial image, containing multiple emotion categories (such as fear, sadness, anger, and happiness). Based on the probability values and a preset threshold, the server determines the dominant emotion category and stores the emotion category and probability as the emotional feature of that frame, along with the frame number and timestamp, as a frame-level emotion record. The server's output is a list of emotion features arranged in a time series.
[0269] Step 8: The server performs pose estimation and action recognition processing and generates action feature quantities.
[0270] The server takes time-series image frames as input, calls a pose estimation algorithm or model, and calculates the position coordinates of key points on the child's body in each frame. Based on the geometric relationships of the key point coordinates and the displacement changes between adjacent frames, the server identifies whether there are defensive postures (e.g., crossed arms, body recoil) and aggressive contact (e.g., someone else's hand suddenly approaches and touches the child's body). The server encodes the identification results into behavior labels, such as "defensive_posture" and "pushed," along with confidence scores. The server stores these action labels and key point features together as frame-level action features. The server's output is a list of action features containing action labels and related parameters for each frame.
[0271] Step 9: The server aggregates emotional and behavioral characteristics within a time window to generate structured statistical data.
[0272] The server takes frame-level emotion feature lists and action feature lists as input, dividing the time into windows according to a predetermined time unit (e.g., the most recent 30 seconds). Within each time window, the server calculates metrics such as the percentage of fear, the total duration of defensive postures, and the number of aggressive contact behaviors. The server normalizes and quantifies these statistical results, for example, expressing the duration in seconds and converting the percentage into a ratio value between 0 and 1. The server organizes the statistical metrics into key-value pairs, forming structured data objects, and associates them with the start and end times of the time window and the device identifier. The server outputs a series of structured statistical data entries at the time window level.
[0273] Step 10: The server generates scene description text and constructs prompt statements.
[0274] The server takes structured statistical data as input, calls the text generation template module, fills each statistical field into a predefined sentence pattern, and generates a scene description text, such as "In the last 30 seconds, the child showed a fear expression 70% of the time; the child maintained a defensive posture for 20 seconds; the system detected 6 suspected aggressive contact behaviors." The server adds task instructions and output format requirements to this, constructing complete prompt statements, such as: "You are a generative AI model specifically designed to analyze children's emotions and behaviors. Based on the statistical characteristics and scenario description given below, please determine whether the child is likely to be bullied, and provide a risk score from 0 to 1 along with a brief explanation. Scenario description: In the last 30 seconds, the child displayed a fearful expression for 70% of the time; the child maintained a defensive posture for 20 seconds; the system detected 6 instances of suspected aggressive contact. Please output the results, including bullying_risk (numerical value) and explanation (textual description)." The server outputs complete prompt text and scene description text for each time window.
[0275] Step 11: The server calls the generative artificial intelligence model and obtains the inference results.
[0276] The server takes the prompt text and scene description text as input, constructs a model request, and sends it as text input to a generative AI model service deployed locally or in the cloud. The server transmits the request to the model inference engine via a network interface and awaits a response. Upon receiving the response, the server parses the numerical bullying risk indicator (bullying_risk) and explanatory statement from the model's output text, and stores them in association with the corresponding time window and device identifier. The server's output is a set of inference results containing a risk score and explanation.
[0277] Step 12: The server determines the warning level based on the reasoning results and generates an evaluation result.
[0278] The server takes the bullying risk indicator output by a generative artificial intelligence model as input, compares this value with multiple preset thresholds (e.g., 0.3, 0.7) to classify different warning levels, such as "Normal," "Caution," "Warning," and "Severe." The server also appends the model's explanation to the assessment record. The server combines time window information, risk score, warning level, and explanation to form a standardized assessment result object. The server writes this assessment result to the database and marks the analysis status of that time window as "Completed." The server's output is assessment result data that can be used for notification and display.
[0279] Step 13: The server sends warning messages and countermeasures to the terminal.
[0280] The server takes the assessment results as input and determines whether the warning level meets the predetermined alarm threshold (e.g., "Warning" or higher). When the conditions are met, the server generates a warning message, including the risk level, risk score, reasoning, and time window information. Based on the assessment results, the server can further generate countermeasure information, such as "It is recommended that the teacher immediately go to the classroom to observe" or "It is recommended to communicate with the child individually." The server packages the warning message and countermeasure information into response data and sends it to the corresponding terminal via the network interface. The server's output is an alarm response message sent to the terminal.
[0281] Step 14: The terminal receives and displays a warning message, triggering a user notification.
[0282] The terminal takes the warning and countermeasure information returned by the server as input, and parses the risk level, time period, and text description. Based on the risk level, the terminal selects an appropriate display style, such as marking high-risk events in red or orange on the interface. The terminal calls the operating system's notification service to generate a local notification message, such as displaying "A child has been detected to have a high risk of being bullied; please check immediately" in the status bar. The terminal presents detailed content in the application interface, including the time window, risk score, model explanation text, and suggested countermeasures. The terminal's output is a user-visual graphical interface and operation entry point.
[0283] Step 15: Users can view the analysis results and provide feedback.
[0284] Users input the notifications and detailed analysis pages displayed on the terminal, and enter the application interface by clicking the notification to view video playback and text descriptions for the corresponding time period. Users determine the validity of the alarm within the interface and can select buttons to "confirm processing" or "mark as false alarm." User feedback is recorded by the terminal and sent back to the server via the network. User output consists of feedback commands and tag information, used for subsequent system adaptive optimization.
[0285] Step 16: The server records historical information and updates processing parameters or model configurations.
[0286] The server takes user feedback, previous inference results, evaluation results, and warning levels as input, and stores them in the history management module. The server statistically analyzes the proportion of false alarms, missed alarms, and valid alarms within a certain time range, and adjusts the confidence threshold, time window length, and emphasized features in the warning statements of the visual model based on these statistics. When needed, the server triggers an offline training process, updating some model weights or warning statement generation rules using accumulated labeled samples. The server's output includes updated model configuration parameters, threshold settings, and warning statement generation strategies, thereby achieving higher detection accuracy and more reasonable alarm behavior in subsequent processing cycles.
[0287] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.
[0288] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."
[0289] Most existing school bullying detection technologies rely on single-modal data processing, such as facial expression recognition based solely on video images or keyword filtering based solely on dialogue text. This makes it difficult to detect real bullying risks in a timely and accurate manner from the complex and ever-changing school environment. Specifically, traditional systems suffer from the following technical problems: First, the server side lacks a unified data structure and temporal-spatial alignment mechanism for visual and auditory information from cameras and microphones, making it impossible to reliably integrate multi-source data within fine-grained time windows, thus reducing the robustness of bullying scene recognition. Second, servers typically use rule matching or simple classification models to independently determine facial features, behavioral features, and verbal aggression, lacking a high-quality prompt statement construction mechanism for generative artificial intelligence models. This prevents the full utilization of the contextual reasoning and scene understanding capabilities of large language models, resulting in insufficient accuracy and interpretability of bullying probability assessment results. Third, when generating alarm messages for guardians and education personnel, servers often use fixed templates or simple filling methods, failing to adaptively adjust the expression of warning content and the granularity of countermeasure suggestions according to different recipient attributes, resulting in insufficient readability and executability of alarm information. Fourth, traditional systems lack a closed-loop learning mechanism based on user feedback and actual event judgment results, making it difficult to update judgment parameters and prompt statement generation strategies in a timely manner, thus limiting the system's adaptive capabilities and detection accuracy improvement in long-term operation.
[0290] This invention aims to provide a server-side processing technology that utilizes multimodal data integration, generative artificial intelligence models, and feedback-driven learning control. By introducing a unified event information generation mechanism for visual and auditory information, an automatic prompt statement generation mechanism for generative artificial intelligence models, and a feedback-based parameter update mechanism into the server, the accuracy, real-time performance, and interpretability of bullying detection and alarm generation are improved from the perspective of computer system architecture and data processing flow, thereby achieving efficient and intelligent monitoring of the risk of school bullying.
[0291] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.
[0292] In this invention, the server includes: a unit for acquiring visual and auditory information from an observation device, and for associating the visual and auditory information with time and location information for acquisition and recording; a unit for performing image processing on the visual information to extract facial expression features and behavioral features, and for calculating emotion and behavioral indicators according to time intervals; a unit for performing acoustic processing and speech recognition on the auditory information to obtain language information, and for calculating aggression and emotion indicators using natural language processing; a unit for integrating visual analysis results with auditory analysis results based on the facial expression features, behavioral features, and language information through the time and location information to generate event information containing multiple types of information; and a unit for... The system automatically generates prompt statements from the event information and previous event information, inputting them into a generative artificial intelligence model, and generates query information containing these prompt statements. It also includes a unit for inputting the query information into the generative artificial intelligence model to obtain an analysis result containing a bullying probability assessment value and its rationale, and for determining the bullying probability based on the assessment value. Furthermore, it includes a unit for generating warning messages in natural language based on the analysis result and the event information, targeting guardians and educational personnel and adjusting the expression according to the recipient's attributes. Finally, it includes a learning control unit for storing and visualizing the warning messages and the event information, receiving response results, and updating the judgment parameters and prompt statement generation parameters accordingly. This allows for an integrated event modeling and reasoning process for multimodal data within the server. By automatically constructing high-quality prompt statements, it fully utilizes the reasoning capabilities of the generative artificial intelligence model and incorporates user feedback to achieve adaptive parameter updates. This improves the accuracy, stability, and interpretability of bullying probability detection at the computer system level, and enhances the adaptability and practicality of alarm message generation and sending processing in different usage scenarios.
[0293] "Observation device" refers to hardware equipment used to acquire visual and / or auditory information in a target environment, including but not limited to cameras, image sensors, microphones, microphone arrays, and data acquisition modules connected to them.
[0294] "Visual information" refers to image or video data collected by observation devices that can represent the appearance and motion state of objects in a scene, including single-frame image sequences, compressed video streams, and their associated time and location information.
[0295] "Audio information" refers to audio data collected by observation devices that can represent the sound content and acoustic characteristics of a scene, including continuous audio streams or segmented audio segments and their associated time and location information.
[0296] "Time information" refers to time-stamped data used to identify the time when visual and auditory information is acquired, including timestamps, time interval identifiers, and time series information used for multi-source data synchronization and alignment.
[0297] "Location information" refers to spatial marker data used to identify the location of visual and auditory information, including pre-defined location numbers, area identifiers, equipment installation location identifiers, and scene identifiers derived from them.
[0298] "Image processing" refers to the preprocessing and feature extraction operations performed by the server on visual information, including color space conversion, image enhancement, object detection, object tracking, and algorithmic processing for recognizing faces or body poses.
[0299] "Expression features" refer to the feature data extracted from a person's facial region through image processing to represent a person's emotional state, including the location of facial key points, local texture features, and the emotion category and its probability calculated based on the model.
[0300] "Behavioral features" refer to the feature data extracted from the posture changes and movement trajectories of people or groups through image processing and time series analysis, which are used to represent their action patterns. These features include behavioral labels such as approaching, moving away, pushing, waving, and surrounding, as well as their confidence levels.
[0301] "Emotion index" refers to a quantitative measure of emotion calculated within a given time interval based on the analysis of facial features and / or auditory information. It includes numerical values representing the intensity of emotions such as anger, fear, sadness, happiness, and neutrality.
[0302] "Behavioral metrics" refer to quantitative data calculated based on behavioral characteristics within a given time interval, used to measure the frequency, duration, or intensity of a specific behavioral pattern.
[0303] "Acoustic processing" refers to the signal processing operations performed by the server on auditory information, including resampling, noise reduction, framing, feature extraction, and format conversion processing in preparation for speech recognition.
[0304] "Speech recognition" refers to the process by which a server uses speech recognition algorithms or external speech recognition services to convert the speech content in auditory information into text.
[0305] "Language information" refers to the dialogue content represented in text form obtained from auditory information through speech recognition, including word sequences, sentences, and corresponding time boundary information.
[0306] Natural Language Processing (NLP) refers to the text analysis operations performed by a server on language information, including word segmentation, part-of-speech tagging, syntactic analysis, sentiment analysis, keyword recognition, and intent classification.
[0307] "Aggression index" refers to a metric obtained by quantifying insulting words, threatening expressions, and imperative statements in language information based on natural language processing results. It is used to represent the degree of language aggression.
[0308] "Visual analysis results" refers to the set of results obtained by the server after analyzing visual information through image processing, including facial expression features, behavioral features, emotion indicators, behavioral indicators, and their corresponding time and location information.
[0309] "Audio analysis results" refers to the set of results obtained by the server after analyzing auditory information through acoustic processing, speech recognition, and natural language processing. These results include language information, aggression indicators, emotion indicators, and their corresponding time and location information.
[0310] "Event information" refers to a structured data unit generated by the server after integrating visual and auditory analysis results based on time and location information. This data unit is used to describe a scene or behavior that occurs in a specific time interval and location and is associated with multimodal information.
[0311] "Past event information" refers to event information that the system has generated and stored in the past, which is used as historical reference data for current event analysis and model reasoning.
[0312] "Generative AI models" refer to AI models that can automatically generate text output based on input text or structured data, including large-scale language models based on deep learning, used for reasoning about scenarios, generating explanations, and generating suggested content.
[0313] "Prompt statements" refer to the input text automatically constructed by the server to guide the generative artificial intelligence model to perform a specific task. This text includes task descriptions, contextual information, constraints, and output format requirements.
[0314] "Query information" refers to the request data sent by the server to the generative artificial intelligence model, which includes at least prompt statements and optional structured supplementary information to trigger the model to perform inference and generate output.
[0315] "Bullying probability rating" refers to a numerical score or level obtained by a generative artificial intelligence model based on event information, used to measure the degree of likelihood of bullying behavior in the event.
[0316] "Analysis results" refers to the output returned by the generative artificial intelligence model based on the query information, including the bullying probability assessment value, corresponding explanation, and optional classification labels or suggestions.
[0317] "Warning messages" refer to notifications written in natural language and generated by the server based on the parsing results and event information. They are used to inform guardians or education-related personnel of potential bullying risks and related background information.
[0318] "Specific countermeasures" refers to the suggested action plans included in the warning message in response to the detected possibility of bullying, including communication suggestions, intervention measures, and key points for follow-up observation.
[0319] "Receiver attributes" refer to user characteristic information related to the recipient of the warning message, including role type (such as guardian, teacher, administrator), permission level, preference settings, etc., which are used to adjust the expression and information granularity of the warning message.
[0320] "External information terminal" refers to user-side devices used to receive and display warning messages, including mobile terminals, fixed terminals, or wearable terminals, which can interact with servers through communication networks.
[0321] The "Display Control Unit" refers to the functional module in the server used to organize and visualize event information, analysis results, and warning messages, and provide them to guardians and education personnel for browsing and operation through a user interface.
[0322] "Response Results" refers to the information provided to the system by guardians or education-related personnel after viewing warning messages and event information, including implemented measures, investigation conclusions, or confirmation results.
[0323] The “Learning Control Unit” is a functional module used to update the parameters of the bullying judgment processing and prompt statement generation processing based on the correspondence between the response results and the bullying probability evaluation value and the actual event judgment results, thereby realizing the adaptive optimization of the detection model and prompt strategy.
[0324] The embodiments of this invention will be described using a school bullying detection system as an example, but this invention is not limited to school settings and can also be applied to environments requiring monitoring of interpersonal conflicts, such as nursing homes, workplaces, and public places. In the following description, the subject is limited to one of the following: server, terminal, or user.
[0325] In one embodiment, the server includes at least one central processing unit (CPU), memory, a network interface, and a communication interface for connecting to external observation devices (including cameras and microphones). The server runs a general-purpose operating system (e.g., Linux) and middleware software (e.g., a web server, a database management system), and installs image processing software libraries (e.g., OpenCV), speech recognition software development kits (e.g., SDKs compatible with cloud-based speech recognition services), natural language processing libraries, and client libraries for accessing generative artificial intelligence models (e.g., HTTP clients compatible with large language model services). The server stores program instructions and various data structures in its memory. When executed by the CPU, these program instructions perform technical processing such as acquiring and storing visual and auditory information, feature extraction, multimodal integration, constructing prompt statements, calling generative artificial intelligence models, and generating alarm messages.
[0326] During system deployment, users install the aforementioned software components on the server and configure the list of monitoring devices, monitoring area identifiers, time synchronization strategies, data sampling parameters, and alarm thresholds through the management interface. In the configuration interface, users associate each camera and microphone with specific location and position identifiers and set parameters such as data sampling frequency, video resolution, and audio sampling rate. Users also configure the server with the access key, request rate limit, and model version number required for generative artificial intelligence model calls to ensure stable and controllable access to external model services.
[0327] In one implementation, the server uses a relational database management system (e.g., a general-purpose relational database) to manage multiple tables, including but not limited to: a raw video metadata table, a raw audio metadata table, a video parsing result table, an audio parsing result table, an event information table, an evaluation result table, an alarm message table, and a feedback record table. The server defines a structured schema for each table, including primary keys, foreign keys, timestamps, location identifiers, object identifiers, and rating fields. Through this pre-designed data structure, the server can efficiently index, join, and retrieve multi-source data based on time and location information, thereby reducing the computational overhead of multi-table join queries and achieving structured and efficient data management.
[0328] In terms of visual information processing, the server uses image processing software libraries to decode, preprocess, and extract features from the video data acquired from the observation device. The server performs operations such as color space conversion (e.g., from color space to grayscale space), noise filtering, and histogram equalization on video frames to improve the robustness of subsequent face detection and pose estimation. In each frame, the server calls a face detection algorithm based on a convolutional neural network or cascaded classifier to calculate the coordinates of bounding boxes for candidate face regions and maps these boxes to a unified coordinate system. If the server detects a face near the same location in multiple consecutive frames, it assigns a temporary object ID to the face using a simple object tracking algorithm (e.g., Kalman filtering combined with nearest neighbor matching) to construct the person's trajectory.
[0329] The server inputs each detected facial region into an expression recognition model. In one implementation, this model can be a convolutional neural network containing multiple convolutional layers, pooling layers, and fully connected layers. Its output is a probability distribution of multiple emotion categories (e.g., anger, fear, sadness, happiness, neutrality). The server extracts the probability value of each emotion category from this probability distribution as part of the expression feature vector and performs sliding window statistics on these features over time. For example, it calculates the average and maximum probability of anger within a certain time window, as well as the duration of consecutively exceeding a certain threshold. The server further models changes in the person's position and limb key points through human pose estimation or optical flow analysis to identify actions such as approaching, moving away, pushing, waving, or blocking, and quantifies the results as behavioral features and indicators.
[0330] In terms of auditory information processing, the server decodes and preprocesses the audio stream acquired from the observation device. The server resamples the audio to a uniform sampling rate and employs bandpass filtering and noise suppression algorithms to reduce environmental noise interference. Subsequently, the server uses a speech activity detection algorithm to segment the audio into segments containing speech content and calls external speech recognition services for these segments. The server sends audio segments and relevant parameters (such as language type and sampling rate) to the speech recognition service through an application programming interface (API) and receives the returned transcribed text and the time boundaries of each sentence. The server stores the transcribed text in an audio parsing result table and records metadata required for aggressive language detection for each sentence.
[0331] During the natural language processing phase, the server uses word segmentation, part-of-speech tagging, and dependency parsing algorithms to process the transcribed text. The server maintains a vocabulary containing offensive words, threatening statements, and exclusionary expressions. Combining statistical language models or deep learning-based text classification models, it calculates aggression metrics for each sentence, such as the proportion of offensive keywords, sentiment polarity scores, and the frequency of imperative expressions. The server can also classify sentences by sentiment based on their position in the dialogue and additional tone information (if supported by an additional acoustic feature model), obtaining sentiment labels such as anger, sarcasm, and agitation. The server aggregates aggression and sentiment metrics at the time window level to obtain a comprehensive aggression score and sentiment distribution for that time window.
[0332] The server aligns and integrates the visual and auditory analysis results using time and location information. Specifically, the server uses a unified timeline to divide the video and audio analysis results into time intervals of equal length. For each interval, it queries the visual and audio analysis result records corresponding to that time period and location identifier. The server constructs an event information data object for each time interval. This object includes: time range, location identifier, a list of involved object IDs, a summary of each object's facial expression and behavioral characteristics, a summary of the dialogue text within the interval, aggression index, emotion index, and historical statistical information (e.g., the number of times a highly aggressive event occurred at that location or for that object within a certain time window in the past). The server stores this event information in an event information table and assigns it a unique event ID.
[0333] The server automatically constructs prompt statements for the generative artificial intelligence model based on event information and past event information. In one implementation, the server predefines multiple prompt statement templates, each containing a fixed task description paragraph, insertable event description placeholders, and output format instructions. For example, the server can generate prompt statements in the following forms: "You are a school bullying risk assessment assistant. Based on the following monitoring data, please determine whether there is a high risk of school bullying, and provide a risk score from 0 to 100, along with your reasoning:" Time: 15:00 on October 5, 2023 Location: Area A of the playground Child A's expression: An angry expression lasting more than 30 seconds, accompanied by aggressive physical actions such as pushing and shoving.
[0334] Child B's facial expression: obvious fear and withdrawal, retreating repeatedly.
[0335] Dialogue transcript summary: 'You're terrible, you're not allowed to play with us,' 'Get lost,' etc.
[0336] Language sentiment analysis: High frequency of aggressive words and intense tone.
[0337] Historical record: The same child A has experienced three similar incidents in the past two weeks.
[0338] Please provide a bullying risk score and explain your main reasons briefly in natural language. Based on the current event information, the server fills the template with information such as time, location, facial expression summary, behavior summary, dialogue summary, and historical statistics to generate a complete prompt statement, which is then encapsulated as part of the query information.
[0339] In one implementation, the server invokes a generative artificial intelligence model, which may employ a language model based on a multi-layer transformer architecture. During the pre-training phase, this model undergoes unsupervised learning based on a large-scale general corpus, updating its weight parameters by minimizing the cross-entropy loss function of next-word prediction or mask word prediction. In this invention, the server does not need to retrain the model; instead, it guides the model to perform a specific reasoning task by constructing prompts with explicit task descriptions and structured contextual information. The server sends the prompts to the model server via a network interface, receives the model's text response, and parses out the bullying probability assessment and reasoning description. For example, the model might output: "Risk Score: 88. Reason: The same person repeatedly uses aggressive language and physical shoving against the same victim, who exhibits persistent fear and withdrawal, consistent with a persistent bullying pattern." The server internally parses the model output, converting "88" into a numerical field and writing it into the evaluation results table, while retaining the reasoning text provided by the model. The server can locally set multiple risk ranges (e.g., 0–39 for low risk, 40–69 for medium risk, and 70–100 for high risk), mapping the scores to risk levels through simple threshold comparisons for subsequent alarm strategy decisions. This judgment mechanism, which combines multimodal features with the reasoning results of generative artificial intelligence models, leverages the model's contextual reasoning and language understanding capabilities compared to traditional approaches based solely on manual rules or single model outputs. This allows the server to achieve higher accuracy and better interpretability when identifying complex bullying behavior patterns.
[0340] When generating warning messages, the server again utilizes generative artificial intelligence models and prompt statements mechanisms. The server constructs a second type of prompt statement to generate natural language instructions for guardians and education-related personnel. For example, the server can generate the following prompt statement: "Based on the following event information, please generate a formal reminder message to be sent to the student's guardian. The language should be objective and calm, avoiding absolute accusations, and provide three actionable family coping suggestions."
[0341] Event Summary: At 3:00 PM on October 5, 2023, in Area A of the playground, the system detected that student A pushed and shoved student B. Student A maintained an angry expression for an extended period, while student B exhibited clear fear and withdrawal. The conversation repeatedly included insulting language such as "You're terrible, you're not allowed to play with us" and "Get lost."
[0342] Assessment conclusion: The generative AI model gives a bullying risk score of 88 (out of 100), which is considered a high risk.
[0343] Please write a reminder message suitable for sending to guardians via a mobile app, and provide three actionable family suggestions. The server submits the prompt and event information together to the generative AI model and receives the message body generated by the model. The server extracts an appropriate title, body, and suggested items from the generated result, combines them into a warning message data object, and stores it in the alarm message table. The server can also adjust the language requirements of the prompt based on the recipient's attributes (e.g., whether it's a homeroom teacher, grade administrator, or guardian), enabling the model to generate text with different styles and levels of detail, thus achieving adaptive support for different user roles.
[0344] In one embodiment, the terminal can be a smartphone, tablet, or desktop computer. The terminal runs a dedicated application or a browser-based client, responsible for receiving and displaying warning messages from the server. The terminal requests a list of unread alerts for the user from the server via a secure communication protocol, parses the structured data returned by the server into interface elements, and displays the event time, location, risk level, summary information, and detailed description on the screen. The terminal also provides input controls, allowing the user to select processing options after viewing the message, such as "Communicated with the student," "Notified the school," or "Confirmed not to be bullying." The terminal sends the user's input response to the server via a network interface to update the feedback log.
[0345] Based on these feedback records and subsequent manually verified event judgments, the server constructs a correspondence between evaluation values and real labels to improve its internal parameter settings and prompt generation strategies. The server can statistically analyze the distribution of aggression indicators, facial expression indicators, and model scores corresponding to events marked as "false alarms" by users within a certain time period. This allows for adjustments to alarm thresholds and prompt content, such as adding emphasis to prompts about "whether persistent behavior exists" or "whether it has occurred repeatedly," making the generative AI model pay more attention to these characteristics in subsequent reasoning, thereby reducing the false alarm rate. Similarly, the server can aggregate events confirmed to involve serious bullying behavior, analyze their typical characteristic combinations (such as high aggression indicators, high anger indicators, persistent fear expressions from victims, and frequent historical occurrences), and reflect these characteristic patterns in the event information construction logic and prompt templates to improve sensitivity to truly high-risk events.
[0346] From a technical perspective, the server achieves efficient integration of multimodal data within fine-grained time windows through unified temporal and spatial alignment and structured modeling of visual and auditory information. This reduces the overhead of separate processing of multi-source data and manual comparison in traditional systems, improving overall processing speed and resource utilization. The server automatically generates high-quality prompts for generative AI models, compressing complex multimodal features into text representations suitable for language model processing. This fully leverages the model's contextual reasoning capabilities without requiring the construction of a complex custom inference engine on the server side. This unconventional processing flow does not simply automate human judgment; rather, it introduces specific data structures, feature combinations, and prompting strategies within the computer, achieving continuous optimization of system behavior through parameterized learning control. This results in significantly superior accuracy, robustness, and interpretability compared to traditional rule-based systems.
[0347] Furthermore, the server can be further improved in terms of data transmission and storage depending on the scenario. For example, during the video parsing stage, the server can retain only high-level features relevant to the detection task (such as emotion vectors and behavioral labels), while compressing and archiving or periodically deleting the original video frames to reduce long-term storage burden and network bandwidth consumption. The server can also adopt a method of extracting features locally before uploading them in audio processing, transmitting only necessary audio segments and features, further reducing cloud access costs and communication load. These optimizations enable the system of this invention to not only improve algorithm accuracy but also achieve improvements in data management and computational efficiency, thus enhancing the computer technology itself.
[0348] In other implementations, the server can employ different neural network structures as facial expression recognition or text classification models. For example, residual network structures can be used to improve the stability of face recognition under complex lighting conditions, or bidirectional recurrent neural networks can be used to model the dialogue context to more accurately identify implicit threats and long-term exclusionary behaviors. The server can also deploy generative artificial intelligence models on local dedicated hardware, such as using graphics processing units or dedicated acceleration chips, to perform inference computation locally, thereby reducing reliance on external services and lowering network latency.
[0349] Through the combination and variation of the above-mentioned various implementation forms, users can flexibly configure the data flow, module composition, and model invocation method on the server side according to the deployment environment, computing resources, and privacy protection requirements, without departing from the technical concept defined by the claims of this invention. Therefore, this invention realizes a complete technical chain from data acquisition, feature extraction, multimodal integration, prompt statement construction, generative artificial intelligence model inference to feedback-driven learning control, providing a high-precision, highly scalable, and highly interpretable technical implementation solution for computer-based bullying risk detection.
[0350] use Figure 13 The processing flow is explained.
[0351] Step 1: The server receives and stores raw data from the observation device.
[0352] Inputs: video stream from the camera, audio stream from the microphone, and device identifier, location identifier, and timestamp for each observation device.
[0353] The server establishes a continuous connection with each observation device via a network interface, receiving encoded video frames and audio data packets according to a preset sampling frequency. Based on the reception time, the server adds a timestamp and location identifier to each video frame and audio segment, writes the video data to a video file or frame buffer, writes the audio data to an audio file or buffer, and generates raw data records in the database. The record fields include device ID, location ID, time range, file path, and data format. Output: Raw video and audio records with time and location information indexes.
[0354] Step 2: The server preprocesses visual information and performs face detection.
[0355] Input: The original video record stored in step 1 (video file path, time range, location ID).
[0356] The server reads video files within a specified time range from the storage device, decodes them into consecutive frame images, and performs image preprocessing operations on each frame, including color space conversion, histogram equalization, and noise reduction filtering. The server calls face detection algorithms from an image processing library to search for face regions in each frame, calculating the bounding box coordinates and confidence score for each candidate face. The server matches face locations across multiple consecutive frames, assigning faces with similar trajectories to the same temporary object ID to form a person tracking sequence. Output: A set of face detection results arranged chronologically, including timestamps, location IDs, object IDs, bounding box coordinates, and detection confidence scores.
[0357] Step 3: The server performs facial expression recognition on the face and calculates emotion indicators.
[0358] Input: The face detection results and corresponding frame images from step 2.
[0359] The server crops the face region from each frame using bounding boxes, normalizes the cropped image to the size and pixel range required by the expression recognition model, and inputs it into the expression recognition neural network. This network outputs probability distributions for various emotion categories. The server uses this to construct an expression feature vector for each object at each time point, and performs statistical calculations on the expression features of the same object within a preset time window, such as calculating the average probability, maximum probability, and duration exceeding a threshold for emotions like anger and fear. The server writes these statistical results as emotion indicators into the video parsing results table. Output: Records of expression features and emotion indicators organized by object ID and time window.
[0360] Step 4: The server analyzes the character's posture and trajectory to extract behavioral features and indicators.
[0361] Input: The trajectory information of the person in step 2 and the corresponding frame image sequence.
[0362] The server performs a human pose estimation algorithm on the human body region, extracts the coordinates of key joints, and calculates motion vectors by combining coordinate changes between consecutive frames. Based on predefined behavior pattern rules or classification models, the server labels certain motion patterns as behaviors such as approaching, moving away, pushing, and surrounding. Within a given time window, the server counts the frequency, duration, and intensity of each behavior, generating behavior feature vectors and behavior index data. The server associates these behavior indices with the emotion indices of the same object and stores them in a video analysis result table. Output: A video behavior analysis record containing the object ID, time window, behavior label, and its intensity index.
[0363] Step 5: The server preprocesses the auditory information and performs speech recognition.
[0364] Input: The original audio record stored in step 1 (audio file path, time range, location ID).
[0365] The server reads the audio file from the storage device, performs resampling, noise reduction, and normalization on the audio, and uses a speech activity detection algorithm to divide the continuous audio into speech segments. The server packages each segment along with language parameters (such as language and sampling rate), calls an external speech recognition service interface, and sends the audio data to the speech recognition engine over the network. The server receives the returned transcription results, including the corresponding text content and time boundaries. The server timestamps the time range and location ID for each transcribed text record and stores it in the audio parsing result table. Output: A collection of dialogue text records with time and location information.
[0366] Step 6: The server performs natural language processing on the transcribed text, calculating aggression and sentiment metrics.
[0367] Input: The dialogue text record obtained in step 5.
[0368] The server performs word segmentation and part-of-speech tagging on each text record, and detects keywords and phrases using a pre-built list of offensive vocabularies and threatening expression patterns. The server can use a text classification model (such as a deep learning-based sentiment analysis network) to perform sentiment polarity analysis on the entire sentence, outputting sentiment scores such as anger and sarcasm. Based on the proportion of offensive words, sentiment scores, and the proportion of imperative sentence structures, the server uses a weighted calculation to obtain aggression indicators and text sentiment indicators. The server aggregates these indicators at the time window level and writes the results to an audio parsing results table. Output: Records of language aggression indicators and sentiment indicators organized by time window and location ID, along with corresponding dialogue summaries.
[0369] Step 7: The server integrates the visual and auditory analysis results to generate event information.
[0370] Input: Video parsing records from steps 3 and 4, and audio parsing records from step 6.
[0371] The server performs join queries on the video and audio parsing tables according to a unified time window and location ID, selecting records with overlapping times and consistent locations. The server combines the object's facial expression and behavioral indicators within each time window with the corresponding verbal aggression and emotion indicators to form an event data object. The server also retrieves relevant events from the historical database concerning the same object or location over a past period, calculates statistics such as frequency of occurrence, and appends them to the event data object. The server stores the complete event data object in an event information table. Output: Event information records containing time range, location identifier, multimodal features, and historical statistical information.
[0372] Step 8: The server automatically generates prompts for the input generative artificial intelligence model based on event information.
[0373] Input: Event information records generated in step 7 and associated historical event information.
[0374] The server reads fields such as event time, location, object facial expression summary, behavior summary, dialogue summary, aggression index, sentiment index, and historical statistics, and fills these fields into a predefined prompt statement template. The server then concatenates a fixed task description, event description, and output requirements to form a natural language prompt statement. For example, the server generates the following prompt statement text: "You are a school bullying risk assessment assistant. Based on the following monitoring data, please determine whether there is a high risk of school bullying, and provide a risk score from 0 to 100, along with your reasoning:" Time: 15:00 on October 5, 2023 Location: Area A of the playground Child A's expression: An angry expression lasting more than 30 seconds, accompanied by aggressive physical actions such as pushing and shoving.
[0375] Child B's facial expression: obvious fear and withdrawal, retreating repeatedly.
[0376] Dialogue transcript summary: 'You're terrible, you're not allowed to play with us,' 'Get lost,' etc.
[0377] Language sentiment analysis: High frequency of aggressive words and intense tone.
[0378] Historical record: The same child A has experienced three similar incidents in the past two weeks.
[0379] Please provide a bullying risk score and explain your main reasons briefly in natural language. The server encapsulates the prompt statement and necessary auxiliary parameters into a query information object. Output: Query information for generative artificial intelligence models, including the complete prompt statement.
[0380] Step 9: The server invokes a generative artificial intelligence model to assess the likelihood of bullying in the event.
[0381] Input: The query information generated in step 8 (including prompts).
[0382] The server sends a request to the generative AI model service via a network interface, providing the prompt as model input and specifying the model version and inference parameters. The generative AI model, based on its internal multi-layered transformer network structure, encodes the event description in the prompt and performs contextual reasoning, outputting text containing a risk score and analytical justification. The server receives this text response, parses the numerical score (e.g., "88") and the textual justification, and performs formatting standardization, such as converting the score to a numerical type and storing the justification by sentence. The server writes the score and justification to an evaluation result table and associates them with the corresponding event ID. Output: An evaluation result record containing a bullying probability assessment value and justification.
[0383] Step 10: The server generates warning messages for guardians and education-related personnel based on the assessment results.
[0384] Input: Event information from step 7 and evaluation results from step 9.
[0385] The server first determines whether an alert is needed based on a set threshold, for example, triggering an alert when the evaluation value is greater than or equal to a certain number. For events requiring an alert, the server constructs a second type of prompt statement to generate a natural language warning message. For example: "Based on the following event information, please generate a formal reminder message to be sent to the student's guardian. The language should be objective and calm, avoiding absolute accusations, and provide three actionable family coping suggestions."
[0386] Event Summary: At 3:00 PM on October 5, 2023, in Area A of the playground, the system detected that student A pushed and shoved student B. Student A maintained an angry expression for an extended period, while student B exhibited clear fear and withdrawal. The conversation repeatedly included insulting language such as "You're terrible, you're not allowed to play with us" and "Get lost."
[0387] Assessment conclusion: The generative AI model gives a bullying risk score of 88 (out of 100), which is considered a high risk.
[0388] Please write a reminder message suitable for sending to guardians via a mobile app, and provide three actionable family suggestions. The server sends the warning statement to the generative AI model, requesting the generation of the corresponding message body. The server receives the text content output by the model, breaks it down into a title, body, and suggestion list, and combines this with metadata such as event time, location, and rating to form a warning message data object, which is then stored in the alert message table. Output: A record of warning messages containing natural language text and structured metadata.
[0389] Step 11: The server sends a warning message to the terminal, which then displays the message and receives user feedback.
[0390] Input: Alarm message records and recipient information generated in step 10.
[0391] The server queries the alarm message table for records in the "Pending Send" status and determines the sending channel and message version based on the recipient attributes. The server sends the message content in a structured format to the corresponding terminal via push service or communication protocol. Upon receiving the message, the terminal parses the title, body, event time, location, and suggested items locally and displays them to the user. The terminal provides feedback input controls for the user, such as buttons or text boxes, allowing the user to select or fill in responses such as "Communicated with the student," "Contacted the school," or "Confirmed no bullying." The terminal sends the user input back to the server via the network interface. The server associates the received response with the corresponding event ID and records it in the feedback table for subsequent learning control and parameter adjustments. Output: Feedback records on the server side and read / feedback status updates on the terminal side.
[0392] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0393] In existing monitoring and analysis technologies, the detection of aggressive or harmful behaviors (such as bullying) targeting children or other vulnerable individuals primarily relies on single-modal data analysis, such as facial expression recognition based solely on images or emotion recognition based solely on speech. This type of approach suffers from the following technical problems: (1) Multimodal information is not effectively integrated. Existing systems typically process images and audio separately, lacking unified emotional state modeling and behavioral pattern modeling. They cannot integrate multi-source data such as facial expression changes, body movements, voice tone, and language content over time, resulting in insufficient robustness and accuracy of risk assessment.
[0394] (2) Computer systems have difficulty automatically generating interpretable natural language security alerts. Traditional systems mostly operate with simple numerical thresholds or fixed template alarms. Although the computing unit can output risk scores, it cannot use high-level semantic understanding to automatically generate scenario-based explanatory information and specific response plans, making it difficult for guardians or educational participants to understand complex test results in a timely manner and take appropriate measures.
[0395] (3) Existing learning mechanisms do not make sufficient use of user feedback. Many systems only perform offline training or periodic parameter updates after deployment, lacking a refined feedback loop for specific alarm events. They do not strictly correlate the post-event evaluation information of guardians or educational participants with the multimodal feature data at that time, thus failing to make targeted model corrections for false alarms and missed alarms, making it difficult to continuously optimize detection performance with the real use environment.
[0396] (4) Generative AI models are not tightly coupled with the perception layer. Although current generative AI models can generate natural language text, they are usually used as independent tools and are not deeply integrated with perception modules such as image recognition, emotion recognition, and behavior analysis at the structured data level. The construction of prompts lacks a systematic design based on multimodal features and risk indicators, making it difficult to ensure the consistency and traceability between the output text and the underlying perception results.
[0397] Therefore, how to construct a system in a computer system that can: (i) integrate multimodal features of dynamic images, audio, and historical behavioral data; (ii) make high-precision judgments on the probability of aggressive or harmful behavior based on unified risk indicators; (iii) automatically generate natural language descriptions and response plans that strictly correspond to the underlying structured information; (iv) form a continuous self-learning and model update closed loop through user evaluation information; and (v) take the internal data processing flow and model collaboration of the computer as the core improvement points has become a technical issue that urgently needs to be solved in this field.
[0398] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.
[0399] In this invention, the server includes: a device for acquiring facial expression and behavioral information of an object from acquired dynamic image and audio information; a device for analyzing the emotional and behavioral states of the object based on the facial expression and behavioral information and extracting various feature information including time-series historical information; a device for calculating risk indicators indicating the likelihood of aggressive or harmful behavior against the object based on the feature information and determining the degree of danger; a device for generating prompt statements based on the facial expression, behavioral information, emotional state, historical information, and degree of danger to enable a generative artificial intelligence model to generate natural language explanation information and coping solution information, and inputting the prompt statements and structured information into the generative artificial intelligence model to obtain the explanation information and coping solution information; a device for adjusting the content and urgency of the explanation information and coping solution information according to the degree of danger and the emotional state and providing it as warning information to guardians or educational participants; and a device for storing evaluation information obtained from guardians or educational participants in association with the feature information and the degree of danger after providing the warning information, and performing learning processing for analysis and degree of danger determination based on the evaluation information to improve the accuracy of risk indicator calculation. This enables the implementation of a unified emotion and behavior modeling mechanism for multimodal data within a computer. Through risk indicator-driven generative AI model prompts and natural language output, the system can not only accurately detect the likelihood of aggressive or harmful behavior but also automatically generate explanations and countermeasures that correspond one-to-one with the underlying feature data and are directly understandable to humans. Simultaneously, by leveraging feedback from evaluation information to establish a closed-loop self-learning process, the parameters of the recognition and generation models are continuously optimized. This results in substantial improvements to computer technology itself at the levels of algorithm structure, data flow, and model collaboration, enhancing the system's detection accuracy, robustness, and interpretability in real-world complex environments.
[0400] A "system" refers to an overall technological structure consisting of multiple information processing devices, storage devices, and communication devices, used to collect, transmit, analyze, and output results from input data.
[0401] "Device" refers to a hardware unit, software unit, or combination thereof used in a system to perform a specific data processing function, including but not limited to processing circuits, program modules, and their operating environment.
[0402] "Acquired dynamic image information" refers to a sequence of image data that has been continuously acquired from physical space and digitized by an image acquisition unit, and contains time-varying data.
[0403] "Audio information" refers to time-series signal data that represents environmental sounds or speech content, which is acquired from physical space and digitized by a sound acquisition unit.
[0404] "Object" refers to the individual or group entity that is being monitored and analyzed, usually the subject that requires security or status monitoring.
[0405] "Facial expression information" refers to the feature data extracted from an image representing the facial region of an object, used to reflect the state of the facial muscles and emotional expression of the object.
[0406] "Behavioral information" refers to feature data extracted from images representing changes in the body or position of an object, used to reflect the object's movement patterns, posture changes, or relative positional relationships.
[0407] "Emotional state" refers to the representation of the category and intensity of an object's psychological emotions at a specific moment or time period, based on multimodal feature analysis.
[0408] "Behavioral state" refers to the classification result of an object's current action type, interaction method, or behavioral pattern based on behavioral information analysis.
[0409] "Time series historical information" refers to a collection of data acquired and stored continuously at multiple points in time, which shows how an object's past expressions, behaviors, and emotions have changed over time.
[0410] "Feature information" refers to the numerical representation obtained by processing the original dynamic image information and audio information, which is used for subsequent recognition or judgment, including but not limited to facial expression features, behavioral features, speech features and text features.
[0411] "Aggressive or harmful behavior" refers to behavioral patterns that may have a negative impact on the physical and mental health of the target, including but not limited to inappropriate behaviors such as insults, threats, rejection, and physical conflict.
[0412] "Risk indicators" refer to numerical values or probabilities calculated based on characteristic information, used to quantify the likelihood of aggressive or harmful behavior towards an object.
[0413] "Risk level" refers to the result of classifying the size of risk according to risk indicators, and is used to indicate the category information of relatively high or low risk.
[0414] "Structured information" refers to a data set organized according to a predetermined data structure, consisting of elements such as characteristic information, emotional state, behavioral state, time information, and degree of danger.
[0415] "Prompt statements" refer to natural language text or its equivalent representation input into a generative artificial intelligence model, used to instruct the model to generate specific types of output content.
[0416] "Generative artificial intelligence models" refer to artificial intelligence models built on a large amount of training data, used to generate natural language text or other data outputs based on input information.
[0417] "Natural language explanatory information" refers to textual information expressed in natural language and output by generative artificial intelligence models to explain detection results, context, and causes of risks.
[0418] "Response plan information" refers to suggestive text information output by generative artificial intelligence models, expressed in natural language, used to guide guardians or educational participants to take specific response measures.
[0419] "Warning information" refers to notification information delivered to guardians or educational participants when the level of danger meets predetermined conditions, indicating the risk of aggressive or harmful behavior.
[0420] "Guardian" refers to an individual or organization that has guardianship responsibility for the life, education, or safety of the individual.
[0421] "Educational stakeholders" refers to individuals or organizations responsible for the learning, management, or guidance of students within an educational environment.
[0422] "Evaluation information" refers to the data formed by the guardian or educational participant after receiving the warning information, including the confirmation results, processing results, or feedback on the actual situation corresponding to the warning.
[0423] "Learning processing" refers to a data-driven process that uses historical feature information and evaluation information to update parameters or optimize the structure of the identification or generation model in order to improve model performance.
[0424] "Recognition model" refers to a machine learning or deep learning model that outputs category labels, probability values, or other judgment results from input feature information.
[0425] "Generative models" refer to machine learning or deep learning models used to generate new data (including natural language text) based on input information.
[0426] "Performance evaluation" refers to the process of quantitatively evaluating the output results of a recognition model or a generation model on validation data using predetermined evaluation indicators.
[0427] "Model structure" refers to the overall configuration of the composition of each layer within the model, including its type, connection method, and dimensional settings.
[0428] "Model parameters" refer to the weights, biases, and other adjustable values obtained through optimization algorithms during training, which determine the model's output behavior.
[0429] In various embodiments of this invention, the server, client, and user each assume different functional roles, collaboratively achieving multimodal analysis of the object's emotional and behavioral states, calculation of risk indicators, generation of natural language descriptions, and self-learning updates. The system of this invention, through specific data structure design, feature extraction algorithms, neural network structures, and the combination of generative artificial intelligence models and prompt statements, achieves high-precision determination of the likelihood of aggressive or harmful behavior, and substantially improves technical performance within the computer.
[0430] I. Overall System Composition A server can be a computing device equipped with a general-purpose processor, a graphics processor, memory, and a network interface. The server runs an operating system and multiple software components, including an image processing library (such as an open-source image processing library), a machine learning framework (such as a general-purpose deep learning framework), a speech recognition service interface, an emotion recognition module, and a generative artificial intelligence model interface module.
[0431] The endpoint can be a smart mobile terminal, a fixed camera terminal, or an edge computing terminal equipped with a camera and microphone. The endpoint runs a data acquisition program, calls the operating system's camera and microphone drivers, generates dynamic image and audio information, and interacts with the server via an encrypted communication protocol.
[0432] Users can be guardians or educational participants. Users access the graphical user interface provided by the server via a client or other computing terminal to view warning messages and explanatory text, and enter evaluation information.
[0433] II. Program Module Composition and Data Structure 1. Server-side module composition The server includes, but is not limited to, the following functional modules: The server includes an image acquisition and reception module. In this module, the server receives encoded video streams from multiple endpoints, decodes the video streams into consecutive image frames, and adds a timestamp, endpoint identifier, and location identifier to each frame. The server stores this information as a record structure with a primary key. For example, the server can store a single image frame as a data record containing the fields {frame ID, endpoint ID, timestamp, location ID, image matrix}.
[0434] The server includes an audio receiving and speech-to-text module. In this module, the server receives audio data uploaded from end-user devices, normalizes the audio data to a uniform sampling rate and bit depth, calls an external speech recognition service interface to transcribe audio segments into text, and associates the text with the time range of the corresponding audio segment, forming a data structure of {segment ID, text content, start time, end time}.
[0435] The server includes image preprocessing and face detection modules. In this module, the server uses an image processing library to perform resizing, color normalization, and denoising operations on each frame of the image. The server then calls a face detection sub-model based on a convolutional neural network to detect head and shoulder regions appearing in each frame. For each detected face, the server generates a record containing {face ID, frame ID, face rectangle coordinates, and face image matrix}.
[0436] The server includes a pose estimation and behavior feature extraction module. In this module, the server uses a pose estimation model (e.g., a model based on convolutional networks and graph structures) to locate human keypoints in each frame. The server calculates features such as joint angle changes, displacement vectors, and accelerations from the keypoint trajectories to represent the object's behavioral state. The server stores the behavior features as a data structure of {object ID, time window ID, pose feature vector, motion feature vector}.
[0437] The server includes an expression recognition and emotion feature extraction module. In this module, the server inputs a face image matrix into an expression classification model based on a convolutional neural network. This model can be structured as a combination of multiple convolutional layers, pooling layers, batch normalization layers, and fully connected layers. The output layer uses a soft maximum function to obtain the probability distribution of various emotions. The server uses the output emotion probabilities and hidden layer features as expression features, recorded as {object ID, timestamp, emotion probability vector, expression embedding vector}.
[0438] The server includes a module for extracting speech emotion and language aggression features. In this module, the server calculates phonetic features for audio segments, such as fundamental frequency, energy, and Mel-frequency cepstral coefficients, and inputs these features into a speech emotion model based on a recurrent neural network or a temporal convolutional network, outputting probability vectors for emotions such as anger, sadness, and fear. Simultaneously, the server performs text sentiment analysis and aggression detection on the transcribed text, constructing bag-of-words vectors or context embedding vectors, and using a classification model to determine whether it contains insults, threats, or exclusionary language. The server stores the results as {segment ID, speech emotion vector, text sentiment vector, aggression flag}.
[0439] The server includes an emotion state fusion engine. In this module, the server aligns facial expression features, behavioral features, speech features, and text features along the timeline. The server then uses a multimodal neural network to map features from different sources to a unified emotion space. The multimodal network can employ a multi-branch structure, with each branch corresponding to a different modality. Dimension matching is performed through several fully connected layers, followed by feature fusion using attention mechanisms or weighted summation in the fusion layer. The server outputs an emotion state vector containing emotions such as anger, sadness, fear, and anxiety.
[0440] The server includes a risk indicator calculation module. In this module, the server combines emotion vectors, behavior vectors, historical behavioral features, and environmental information into a high-dimensional input vector, which is then fed into a supervised learning model. This learning model can be a multilayer perceptron, a time series model, or a graph neural network, outputting one or more continuous values representing the risk indicator of the likelihood of aggressive or harmful behavior. Based on the risk indicator and pre-set thresholds, the server classifies the risk into multiple hazard levels.
[0441] The server includes a prompt generation module. Within this module, the server generates natural language prompts based on structured information (including time, location, object identifier, main emotion tags, main behavioral descriptions, historical statistical indicators, and risk level). The server fills the structured fields into a template through template filling and rule concatenation to obtain detailed input text. These prompts drive the generative AI model to generate explanatory and coping information.
[0442] The server includes a generative AI model interface module. In this module, the server sends the prompt along with some structured background information to the generative AI model's interface. The generative AI model can be a large-scale language model based on an attention mechanism, internally employing multi-layered self-attention networks, feedforward neural networks, and positional embedding structures. After receiving the natural language text generated by the model, the server parses the text into two parts: explanatory information and response suggestions.
[0443] The server includes a warning message generation and output module. Within this module, the server adjusts the wording and urgency of the explanatory and response information based on the level of danger and emotional intensity. The server then encapsulates the final warning message into a structure containing a title, summary, detailed description, recommended actions, and relevant time and location identifiers, and sends it to the client or other user terminals.
[0444] The server includes an evaluation information receiving and learning module. In this module, the server receives evaluation information input by the user, such as "confirmed as aggressive behavior," "false alarm," or "abnormal emotion but not aggressive." The server associates the evaluation information with the current feature information and risk indicator records to form new labeled samples, which are then stored in the training dataset. The server performs batch learning processing at predetermined intervals, updating the parameters of the recognition and generation models using loss functions (such as cross-entropy loss or mean squared error). The server updates the model weights using gradient descent-like optimization algorithms, thereby gradually improving the accuracy and rationality of risk indicator calculation and prompt statement generation.
[0445] 2. End-to-end module configuration The endpoint includes a sensor control and data acquisition module. Within this module, the endpoint periodically triggers camera sampling to generate video frames and controls the microphone to sample the audio stream. The endpoint can adjust the sampling frame rate, resolution, and audio sampling rate based on configurations issued by the server.
[0446] The endpoint includes a local preprocessing and compression module. Within this module, the endpoint performs image resizing, basic noise reduction, and lightweight face or motion detection as needed. If the endpoint does not detect a face or motion, it may not upload the corresponding frame to reduce the amount of uplink data. The endpoint encodes and compresses the video and audio to form data packets.
[0447] The endpoint includes a data transmission module. Within this module, the endpoint uses an encrypted transmission protocol to establish a connection with the server and sends the sampled data and its metadata to the server according to a predetermined data format.
[0448] The endpoint includes an alarm display and interaction module. In this module, the endpoint receives warning information pushed by the server and displays it on the screen in a graphical interface. The endpoint can display the risk level, explanatory text, and suggested measures. It can also provide buttons or input boxes for users to submit evaluation information and remarks.
[0449] III. Detailed Explanation and Technical Implementation of the Program Processing When implementing the above modules, the server can deploy each functional module as an independent service process or microservice. The server can use message queues or asynchronous task systems internally to separate image preprocessing, feature extraction, risk calculation, and generative text generation for execution, thereby making full use of multi-core processor and graphics processor resources.
[0450] In the facial expression recognition model, the server can employ a convolutional neural network. Its convolutional layers use several 3×3 convolutional kernels with a stride of 1, and utilize activation functions for non-linear transformations. Pooling layers can use max pooling to compress spatial dimensions. Finally, the fully connected layer outputs the emotion category probability, and the server performs soft maximum normalization on the output. During the training phase, the server uses a classification loss function and employs stochastic gradient descent or its variants for weight updates.
[0451] In its risk indicator calculation model, the server utilizes a fully connected network with several hidden layers to fuse multimodal features. At the input layer, the server concatenates feature vectors from different modalities in a predetermined order and then normalizes them. The server employs activation functions in the hidden layers to introduce non-linearity. The output layer uses either a linear or sigmoid function to output a continuous risk value. During training, the server uses labels with both positive and negative samples to minimize the error between the predicted risk value and the true label.
[0452] In the prompt message generation section, the server defines prompt message templates using rules, for example: "You are an expert in campus safety and child psychology."
[0453] Event Information: - Time: [Time field] - Location: [Location field] - Behavior: [Behavior description field] - Emoticons: [Emoticon tags and intensity fields] - Language: [Offensive Language Summary Field] - History: [Historical behavior and sentiment statistics fields] Please determine whether there is any high-risk aggressive or harmful behavior, and provide a reason in no more than a few words, along with several specific and actionable countermeasures for the responsible parties. The server replaces the placeholders in the template with the corresponding fields in the structured information to obtain the complete prompt statement. The server then sends this prompt statement as input text to the generative AI model interface. Internally, the generative model uses a multi-layered self-attention mechanism to encode the prompt statement and generate response text word-by-word or sub-word-by-sub-word based on probability distributions. After receiving the response text, the server divides it into a risk description section and a response solution section, storing and displaying them as explanatory and response information, respectively.
[0454] In another specific example, the server can use the following prompt statement: "Please write a short suggestion, no more than a few words, for the guardian, in the capacity of a child psychologist."
[0455] Situation description: - The child has recently been showing 'sadness' or 'no obvious expression' at home; - Speaks at a low volume and slowly when talking to family members; School surveillance data shows that he is mostly alone during breaks.
[0456] Please use a gentle, non-accusatory tone to help the guardian understand the child's possible psychological state, and provide several communication and support methods that can be tried immediately. Through the design of the above prompts, the server can guide the generative artificial intelligence model to output natural language descriptions that are highly consistent with the underlying structured data, rather than simply applying a fixed template.
[0457] IV. Specific Application Examples and Technical Effects In one embodiment, the endpoint is installed in a school corridor to continuously capture images of student activities. The server, through image preprocessing and face detection modules, identifies an individual who is cornered by other individuals at multiple time periods. The server detects multiple pushing actions through pose estimation and behavioral feature modules, and detects that the individual is experiencing prolonged fear and sadness through an expression recognition module. The server integrates these features into a risk index calculation module, inputting them into a model to derive a high-risk index value. The server then constructs a warning message, calls a generative AI model to generate a risk description and response plan, and pushes the warning information to the endpoint. After viewing the warning on the endpoint, the user goes to the site to confirm. If aggressive behavior is confirmed, the user selects "Confirmed as Aggressive Behavior" on the endpoint interface. This evaluation information is then transmitted back to the server to update the model.
[0458] Through the aforementioned linkage, the server not only reduces the burden of manual monitoring in terms of quantity, but also achieves automatic fusion of multimodal features and dynamic threshold adjustment in terms of computing structure. The server continuously optimizes parameters through a self-learning process, enabling the system to maintain a high detection rate and a low false alarm rate when facing changing environments and behavioral patterns.
[0459] V. Explanation of Technological Improvements and Causal Relationships The server improves computer technology itself in the following ways by using the multimodal fusion structure and generative artificial intelligence model-driven description generation mechanism proposed in this invention: By using specific data structures internally, the server indexes images, audio, text, and emotion results uniformly according to a timeline, achieving efficient data retrieval and caching, reducing redundant calculations, and improving memory usage efficiency.
[0460] By inputting multimodal features into a unified deep learning model and jointly optimizing the loss function, the server enables the model to output risk indicators in a single forward inference process, reducing the overhead of traditional serial rule comparison and independent inference of multiple models, thereby improving processing speed.
[0461] The server incorporates real-world feedback into its training data through a self-learning process based on user reviews. This allows the model to automatically suppress frequent false positive patterns in specific environments, reducing the overall error rate. This process is achieved through explicit loss functions and weight update algorithms, rather than simple manual rule adjustments.
[0462] By employing a structured design for its prompts, the server ensures a strict correspondence between the output of the generative AI model and its underlying features. This avoids the disconnect between the output of traditional generative models and actual monitoring data, thereby improving the interpretability and credibility of the natural language descriptions. This structured prompt guidance differs from simply writing reports by humans; it represents a purposeful guidance of the probability distribution within the language model.
[0463] By performing lightweight face or motion detection locally and filtering uploaded frames, the endpoint reduces network bandwidth usage and server load, thereby improving overall system communication efficiency and computing resource allocation.
[0464] VI. Alternative Methods and Variations Servers can use different neural network structures in different implementations. For example, a server can use a convolutional-recurrent hybrid network instead of a pure convolutional network, a temporal convolutional network instead of a recurrent structure, or a graph neural network to model the interaction relationships between multiple objects. Servers can employ different strategies for multimodal fusion, such as weighted summation, concatenated fully connected layers, attention fusion, or gating mechanisms, to adapt to different scenarios.
[0465] In terms of hardware selection, the terminal can use cameras with different resolutions and frame rates, as well as infrared or depth cameras, to enhance recognition capabilities in low-light environments. The user terminal can be a mobile phone, tablet, or desktop computer, accessing the server using a browser or dedicated application.
[0466] The server can be deployed locally or remotely for generative AI models, and the language, length, and style of the prompts can be adjusted according to the application scenario. In some implementations, the server can perform rule-based post-processing or confidence checks on the text returned by the generative model, such as filtering inappropriate language or limiting the output length.
[0467] By combining the above-mentioned various implementation forms, users can choose the appropriate configuration according to their actual environment and resource conditions. The core of this invention lies in the server's unified modeling of multimodal features and risk indicators, the invocation of generative artificial intelligence models driven by prompt statements, and the closed-loop self-learning based on evaluation information. These technical points can be implemented under different hardware and software configurations, thus providing a scalable, high-precision, and adaptive computer implementation scheme for the intelligent monitoring of aggressive or harmful behaviors.
[0468] use Figure 14 The processing flow is explained.
[0469] Step 1: The device collects raw data from each end and sends it to the server.
[0470] The endpoint uses a camera and microphone to capture dynamic image and audio information containing objects as input. It performs scaling, frame rate control, and noise suppression on the input video frames, and performs sampling rate unification and simple noise reduction on the audio stream, outputting encoded video and audio data packets. The endpoint discards irrelevant frames based on whether it detects a face or significant movement, and uploads the retained data packets, along with corresponding timestamps, terminal identifiers, location identifiers, and other metadata, to the server via an encrypted communication protocol.
[0471] Step 2: The server receives and preprocesses the image data.
[0472] The server takes the video data packets and their metadata uploaded from each end as input, uses a decoding library to decode the video data into consecutive image frames, and outputs a time-ordered frame sequence. The server performs image preprocessing operations on each frame, including resizing, brightness and contrast normalization, and filtering for noise reduction. The processed frame, along with its timestamp, terminal ID, and location ID, is stored as a structured record. The server thus completes the conversion from compressed video data to a standardized frame matrix.
[0473] Step 3: The server performs face detection and face region cropping.
[0474] The server takes preprocessed image frames as input, calls a face detection model based on convolutional neural networks or other detection algorithms, scans each frame, and detects all face bounding boxes. The server then crops the detected face regions from the entire frame, assigns a temporary object ID to each region, and outputs a set of {temporary object ID, face image matrix, timestamp, location ID}. Through this data processing, the server converts the original panoramic image into local facial images for each object, facilitating subsequent expression recognition.
[0475] Step 4: The server performs pose estimation and behavioral feature extraction.
[0476] The server takes a preprocessed full-frame image and a temporary object ID as input, calls a pose estimation algorithm, and performs keypoint detection on the image region containing the object to obtain the coordinates of the head, torso, and limb joints. The server aligns the keypoint coordinates in multiple consecutive frames by time, calculates displacement vectors, velocities, and accelerations, and outputs a set of behavioral feature vectors corresponding to each object. Through numerical calculations on the keypoint time series, the server obtains quantitative features representing behavioral patterns such as "pushing, surrounding, chasing, stillness, and isolation."
[0477] Step 5: The server performs facial expression recognition and generates emotional features.
[0478] The server takes the face image matrix obtained in step 3 as input, calls the expression classification model based on a convolutional neural network, performs forward inference on each image, and outputs a multi-dimensional emotion probability vector and an intermediate layer embedding vector. The server combines these vectors with the object's temporary ID and timestamp to form a structured output of {object ID, timestamp, emotion probability vector, expression embedding}. The server performs convolution operations, activation function transformations, and fully connected mappings on the pixel matrix to achieve a non-linear mapping from the original image to the emotion feature space.
[0479] Step 6: The server performs audio preprocessing and speech transcription.
[0480] The server takes the audio data packets uploaded from each client as input, performs sampling rate resampling, bandpass filtering, and noise suppression on the audio data to obtain standardized audio segments. The server calls the speech recognition service interface to transmit the audio segments to the speech recognition engine and receives the transcribed text results. The server associates the audio segments with corresponding text, time ranges, and objects, outputting a data record of {segment ID, text content, start time, end time, object ID}. Through waveform signal processing and calls to external recognition services, the server completes the conversion from audio to text.
[0481] Step 7: The server extracts voice emotion and language attack characteristics.
[0482] The server takes standardized audio segments and transcribed text as input. First, it calculates acoustic features such as Mel-frequency cepstral coefficients, fundamental frequency, energy, and speech rate for the audio. Then, it inputs these feature vectors into a speech emotion recognition model to obtain the probabilities of emotions such as anger, sadness, and fear. Simultaneously, the server segments and vectorizes the transcribed text, embedding the text into a text sentiment classification model and an offensive language recognition model to determine whether it contains insults, threats, or exclusionary words. The server outputs a set of {segment ID, speech emotion probability vector, text sentiment label, and offensive language marker}. Through numerical feature extraction and classification operations, the server converts audio and text information into emotion and aggression features that can be used for risk assessment.
[0483] Step 8: The server performs multimodal emotion state fusion.
[0484] The server takes facial expression features, behavioral features, voice emotion features, and text attack features as input. It performs time alignment and interpolation on all features based on timestamps and object IDs to construct a unified multimodal input vector. This vector is then fed into a multimodal fusion network. Each modality undergoes linear transformations and nonlinear activations in its respective subnetwork, and the resulting vectors are integrated at the fusion layer using attention mechanisms or weighted summation. The output is a unified emotion state vector. The server records the result as {object ID, time window ID, emotion state vector}, thus achieving the transformation from scattered features to a unified emotion representation.
[0485] Step 9: The server calculates risk indicators and determines the level of danger.
[0486] The server takes emotional state vectors, behavioral feature vectors, and time-series historical information as input, concatenates them into a high-dimensional input vector, and feeds it into a risk assessment model. Within this model, the server calculates one or more continuous risk values using multi-layer fully connected layers and activation functions, representing the probability of aggressive or harmful behavior occurring. The server compares the output risk value with a preset threshold, classifying it into different risk levels (e.g., low, medium, high) based on the comparison result. The output is {object ID, time window ID, risk value, risk level}. Through this numerical calculation and threshold determination, the server compresses complex multimodal data into risk indicators that are easy to use for subsequent decision-making.
[0487] Step 10: The server generates structured event information.
[0488] The server takes risk records that meet preset danger levels and their corresponding feature data as input, and aggregates them to generate structured event objects. The server extracts key fields from emotional states, behavioral tags, language attack characteristics, and historical statistics, combining them into a structured output of {Event ID, Object ID, Time Range, Location, Main Emotional Tags, Main Behavioral Descriptions, Language Attack Summary, Risk Value, Danger Level}. Through field extraction and assembly operations, the server organizes the underlying feature data into high-level semantic information that can be used to construct prompt statements.
[0489] Step 11: The server constructs a prompt statement and invokes a generative artificial intelligence model.
[0490] The server takes a structured event object as input and generates a prompt statement based on a predefined natural language template. The server fills the template placeholders with time, location, behavioral description, emotional intensity, verbal attack summary, and history to form the complete text input. For example, the server can generate the following prompt statement: "You are an expert in campus safety and child psychology."
[0491] Event Information: - Time: [Time field] - Location: [Location field] - Behavior: [Behavior description field] - Emoticons: [Emoticon tags and intensity fields] - Language: [Offensive Language Summary Field] - History: [Historical behavior and sentiment statistics fields] Please determine if there is a high risk of school bullying, and explain your reasoning in no more than 150 words, providing two specific and actionable suggestions for both the teacher and the guardian. The server sends the prompt statement along with necessary structured background information as input to the generative AI model interface, and outputs natural language description text and solution text generated by the model. The server achieves the conversion from structured information to natural language output through text concatenation and API calls.
[0492] Step 12: The server adjusts the instructions and generates a warning message.
[0493] The server takes as input the explanatory text and response plan text returned by the generative AI model, along with the danger level and emotional intensity. It then post-processes the text content based on the danger level. For high-danger levels, the server can add urgent prompts and instructions requiring immediate intervention; for medium-danger levels, the server can soften the tone and emphasize observation and communication. The server combines the processed explanatory text, response plan, relevant screenshots, and metadata into a warning message object, outputting a serializable notification data structure. Through rule-based adjustments and synthesis, the server creates a warning message that is readable by the user and consistent with the underlying data.
[0494] Step 13: The terminal receives and displays warning messages.
[0495] The endpoint takes the warning message pushed by the server as input and extracts fields such as title, summary, detailed description, suggested measures, and time and location through the local parsing module. The endpoint then presents this information in a graphic and textual format on the display interface, outputting a visual alarm interface. The endpoint can be set with different colors or icons according to the level of danger, and can trigger alert actions such as sound and vibration to help users quickly perceive the severity of the event.
[0496] Step 14: Users can view the warning and enter feedback.
[0497] Users take the alarm interface displayed on the client as input, read the explanatory text and response plan, and make a judgment based on the actual situation on site. Users can select options such as "confirmed as aggressive behavior," "false alarm," and "abnormal emotion but not aggressive" on the client interface, and can add text notes as output. The user's operation results are packaged by the client into an evaluation information object, including event ID, evaluation tag, and notes, for feedback to the server.
[0498] Step 15: Evaluation information is uploaded from each device and stored and associated on the server.
[0499] The client takes the user-inputted evaluation information as input and uploads it to the server via encrypted communication. After receiving the evaluation information, the server associates it with the corresponding event ID, feature information, and risk indicator records, outputting an expanded training sample record. The server stores this record in the training data storage area, providing labeled data for subsequent learning processing.
[0500] Step 16: The server performs self-learning to update the recognition model and generation strategy.
[0501] The server takes the newly accumulated training sample set as input and extracts multimodal feature vectors, corresponding ground truth labels (e.g., whether it is a real attack behavior), and evaluation categories. During the training phase, the server performs forward and backward propagation: by calculating the loss between the predicted risk value and the ground truth label (e.g., cross-entropy loss), the server calculates the gradient and uses optimization algorithms to update the network weights, outputting the updated recognition model parameters. Similarly, the server can adjust the prompt statement template or generative model invocation strategy based on user feedback on the reasonableness of the generated text. Through this data-driven parameter update process, the server gradually reduces the false positive and false negative rates, thereby improving the overall system's detection accuracy and stability in complex environments.
[0502] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0503] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0504] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.
[0505] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0506] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.
[0507] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.
[0508] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.
[0509] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0510] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.
[0511] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0512] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0513] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0514] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0515] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0516] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0517] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.
[0518] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0519] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0520] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0521] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0522] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0523] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0524] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0525] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.
[0526] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0527] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.
[0528] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.
[0529] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.
[0530] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0531] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.
[0532] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0533] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0534] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0535] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0536] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0537] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0538] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.
[0539] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".
[0540] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0541] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0542] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0543] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0544] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0545] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0546] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.
[0547] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0548] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.
[0549] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.
[0550] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.
[0551] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0552] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.
[0553] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0554] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).
[0555] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0556] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0557] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0558] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0559] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0560] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.
[0561] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".
[0562] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0563] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0564] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0565] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0566] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0567] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0568] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.
[0569] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0570] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.
[0571] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.
[0572] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.
[0573] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.
[0574] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).
[0575] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.
[0576] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."
[0577] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.
[0578] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).
[0579] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.
[0580] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0581] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.
[0582] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.
[0583] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.
[0584] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.
[0585] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.
[0586] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.
[0587] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.
[0588] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.
[0589] In addition, the following notes are provided in response to the above explanation.
[0590] Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for acquiring image information containing a child’s facial expressions and behavior from an imaging device, and for attaching time information and recognition information when acquiring the image information in real time; A device for compressing and encoding the acquired image information before sending it to the information processing device, and for sending the image information to the information processing device via a secure communication method; An apparatus for decoding received image information in the information processing device, extracting the facial and body regions of a person using image processing technology, performing emotional state estimation and recognition processing on the facial region, and performing posture estimation and action classification on the body region to extract aggressive and vulnerable behaviors. An apparatus for generating a feature quantity that integrates information on temporal changes in emotional state with time-series information on the aggressive and victimized behaviors, using the feature quantity to calculate an assessment value representing the likelihood of bullying, and determining the likelihood of bullying according to predetermined criteria. An apparatus for generating a prompt statement for a generative artificial intelligence model based on the determination result of the bullying probability, taking contextual information and object information containing elements of warning content as input information, generating a prompt statement for the generative artificial intelligence model, and using the prompt statement to cause the generative artificial intelligence model to generate a warning message and text information about the response strategy, and adjusting the expression content and level of detail of the warning message according to the evaluation value and the recipient attributes. A device for sending the warning message and the text information about the response strategy to an information display device used by guardians and education personnel, and for visually presenting the information display device together with the determination of the likelihood of bullying, the emotional state, and a summary of the behavior analysis. And an apparatus for storing the determination result of the bullying probability, the suggestions given by the generative artificial intelligence model, and the response results and evaluation information of the guardian or education-related personnel as learning data, and for using machine learning algorithms to update the parameters of the identification model and evaluation model used for the emotional state estimation, the action classification, and the bullying probability determination, so as to improve the detection accuracy.
[0591] (Note 2) According to the information processing system described in Appendix 1, the information processing device further includes a device for generating a visualization interface, which displays the determination result of the bullying probability, the temporal changes of the emotional state, the action classification result, the evaluation value, and the suggestions given by the generative artificial intelligence model in chronological order and object units, and provides remotely accessible information presentation user display functions to guardians and education-related personnel through the visualization interface.
[0592] (Note 3) According to the information processing system described in Appendix 1, the information processing apparatus further includes a device for periodically re-evaluating the benchmark value and weight used in calculating the evaluation value based on user evaluation information containing false positives and false negatives; a device for calculating performance indicators for the identification model and the evaluation model and determining whether the model needs to be updated; and a device for automatically performing relearning processing using the machine learning algorithm when it is determined that an update is needed, so as to continuously improve the accuracy of bullying probability detection and the adaptability of warning messages.
[0593] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for receiving image information and corresponding attribute information from a terminal that collects children's image information; A device for converting the image information into time-series image information and performing target recognition processing, facial expression recognition processing and motion recognition processing on the image information to extract feature quantities related to the child's emotional state and body movements; A device for summarizing the feature quantities according to a predetermined time unit to generate structured data as statistical information, wherein the statistical information includes at least the proportion of emotional states occurring within the time unit, the duration of defensive postures, and the number of contact behaviors from others. An apparatus for automatically generating prompt statements for inputting into a generative artificial intelligence model based on the structured data and text information for describing the structured data, and inputting the prompt statements and the text information into the generative artificial intelligence model; A device for assessing the likelihood of a child being bullied based on inference results output from the generative artificial intelligence model, the inference results including at least a bullying risk indication expressed in numerical form and a reasoning explanation for the bullying risk indication, the assessment including a device for determining the bullying risk indication as a warning level; An apparatus for sending a warning message containing the assessment results and the explanation of the reasons to the terminal when the warning level meets a predetermined benchmark, and for generating a warning message for notification to the guardian or education-related party based on the warning message; A device for generating, based on the assessment results, information on coping strategies and candidate intervention behaviors for the child.
[0594] (Note 2) The information processing system according to Appendix 1 is characterized in that, The device for receiving image information and corresponding attribute information from a terminal that collects children's image information is configured to extract the feature quantity and generate the structured data in real time or near real time from the image information sent by the terminal, and the system continuously performs the processing of providing the prompt statement to the generative artificial intelligence model, receiving the inference result and sending the warning information, thereby continuously monitoring the possibility of the child being bullied.
[0595] (Note 3) The information processing system according to Appendix 1 is characterized in that, The system is configured to store the inference results, the evaluation results, the warning level, and the response information of the guardian or education-related party corresponding to the warning information as historical information, and to update the target recognition processing, the facial expression recognition processing, the action recognition processing, and the generation conditions of the prompt statements for the generative artificial intelligence model based on the historical information, thereby continuously improving the accuracy of the assessment of the possibility of the child being bullied.
[0596] Example 2 (Note 1) An information processing system, characterized in that it comprises: A unit for acquiring and recording visual and auditory information from an observation device, and for associating the visual and auditory information with time and location information; A unit for performing image processing on the visual information to extract facial expression features and behavioral features of a person, and to calculate emotion indicators and behavioral indicators by time interval; A unit for performing acoustic processing and speech recognition on the auditory information to obtain language information, and for calculating aggression and emotion indices using natural language processing. This unit is used to integrate visual analysis results with auditory analysis results based on the facial expression features, behavioral features, and language information, through the time information and location information, thereby generating an event information unit containing multiple types of information; A unit for automatically generating prompt statements to be input into a generative artificial intelligence model based on the event information and previous event information, and generating query information containing the prompt statements; A unit for inputting the query information into the generative artificial intelligence model to obtain an analysis result containing a bullying probability evaluation value and its reasons, and for determining the bullying probability based on the evaluation value; A unit for generating, in natural language, a warning message containing warning content and specific countermeasures for guardians and education-related personnel based on the parsing results and the event information, and adjusting the expression of the warning content according to the recipient attributes; A unit for sending the warning message to an external information terminal.
[0597] (Note 2) The information processing system according to Appendix 1 is characterized in that it further includes: This system is used to store the warning messages and event information, and to visualize them according to the timeline, location, object attributes, and evaluation values. This allows guardians and education personnel to browse the event information, analysis results, and past response records through the user interface, and to send response results via input.
[0598] (Note 3) The information processing system according to Appendix 1 is characterized in that it further includes: This learning control unit is used to store the response results and the correspondence between the evaluation values and the actual event judgment results, and to update the parameters of the bullying judgment processing and prompt statement generation processing based on the correspondence, thereby adjusting the input form and query content for the generative artificial intelligence model and improving the accuracy of bullying probability detection.
[0599] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for obtaining facial expression and behavioral information of an object from acquired dynamic image and audio information; An apparatus for analyzing the emotional and behavioral states of an object based on the facial expression information and the behavioral information, and extracting various feature information including time-series historical information; A device for calculating a risk index based on the feature information that indicates the likelihood of aggressive or harmful behavior against the object, and for determining the degree of danger based on the risk index; An apparatus for generating prompt statements based on the facial expression information, behavioral information, emotional state, historical information, and degree of danger, enabling a generative artificial intelligence model to generate natural language explanation information and coping solution information, and inputting the prompt statements and the structured information into the generative artificial intelligence model to obtain the explanation information and the coping solution information; A device for adjusting the content and urgency of the explanatory information and the coping strategy information according to the degree of danger and the emotional state, and providing them as warning information to the guardians or educational participants related to the object; An apparatus for storing, after providing the warning information, evaluation information obtained from the guardian or the educational participant in association with the feature information and the degree of danger, and for performing learning processing for the analysis and the determination of the degree of danger based on the evaluation information, so as to improve the accuracy of the risk indicator calculation.
[0600] (Note 2) The information processing system according to Appendix 1 is characterized in that, The system further includes: a device for generating an information display screen, which can display the feature information, the degree of danger, and the evaluation information, including the explanatory information and the response plan information, in a time-series manner, and correspond to multiple monitoring target spaces and multiple objects, so that the guardian or the educational participant can access the information display screen through a remote terminal to browse the warning information and the historical information, and input the evaluation information through an operation interface.
[0601] (Note 3) The information processing system according to Appendix 1 is characterized in that, The learning process includes: evaluating the performance of a recognition model or generation model used to perform at least one of the following processes within a predetermined period, based on a learning information set containing recently stored feature information and evaluation information; and updating the structure or parameters of the model according to the evaluation results, so as to continuously improve the detection accuracy of the probability of the occurrence of the aggressive or harmful behavior and the rationality of the warning information.
Claims
1. An information processing system, characterized in that, include: processor; The processor is configured as follows: Image recognition and emotion analysis technologies are used to analyze children's facial expressions and behaviors to detect the possibility of bullying; Generative artificial intelligence models are used to generate warning messages, and the warning content is adjusted accordingly. Provide guardians and education-related personnel with specific countermeasures based on the analysis results.
2. The information processing system according to claim 1, characterized in that, The processor is configured to provide a dashboard for presenting the results of children’s facial expression and behavior analysis to guardians and education stakeholders, and to enable guardians and education stakeholders to view the information through an interface accessible to them.
3. The information processing system according to claim 1, characterized in that, The processor is configured to learn from newly acquired data and periodically evaluate the model to improve the accuracy of bullying probability detection, thus forming a feedback loop for improving accuracy.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A