ShypdShypd.ai
🤖

AI Agents & Automation

Browsing page 352 of AI Agents & Automation. Sorted by confidence score — our independent quality rating.

DeepLearningExamples

DeepLearningExamples

60%

DeepLearningExamples is a comprehensive repository from NVIDIA, offering state-of-the-art deep learning scripts. These examples are meticulously organized by models, making them easy to train and deploy while ensuring reproducible accuracy and performance. The platform is designed for enterprise-grade infrastructure, leveraging the NVIDIA CUDA-X software stack and optimized for NVIDIA Volta, Turing, and Ampere GPUs. It includes a wide array of models across computer vision, natural language processing, recommender systems, speech to text, text to speech, graph neural networks, and time-series forecasting. The examples are provided within monthly updated Docker containers on the NGC container registry, ensuring users have access to the latest NVIDIA examples, framework contributions, and optimized deep learning software libraries like cuDNN and NCCL.

trajectory-transformer

trajectory-transformer

60%

Trajectory Transformer is an open-source code release that implements offline reinforcement learning as a sequence modeling problem. Based on the paper "Offline Reinforcement Learning as One Big Sequence Modeling Problem," this tool provides a framework for training models to predict trajectories. It includes scripts for training transformers on various datasets and for planning with these models. The project also offers pretrained models for multiple datasets, allowing users to quickly experiment and reproduce results. It supports installation via conda or Docker, and provides utilities for running jobs on Azure, making it suitable for researchers and engineers in reinforcement learning and robotics.

tokenizers

tokenizers

60%

tokenizers is an open-source library developed by Hugging Face, offering highly optimized and versatile tokenizers for natural language processing tasks. Implemented primarily in Rust, it boasts exceptional performance, capable of tokenizing a gigabyte of text on a server's CPU in less than 20 seconds. The library supports training new vocabularies and tokenizing text using popular models like Byte-Pair Encoding, WordPiece, and Unigram. It includes features such as alignment tracking during normalization, ensuring that the original sentence segments corresponding to tokens can always be retrieved. Additionally, it handles pre-processing steps like truncation, padding, and adding special tokens required by various models, making it suitable for both research and production environments.

Tarot Master

Tarot Master

60%

Tarot Master is an innovative platform that combines the mystical wisdom of Tarot with the precise insights of Astrology, enhanced by artificial intelligence. Users can chat with their personal AI psychic to receive highly personalized insights based on their unique astrological data. The platform offers 24/7 availability with over 25 AI-enhanced Tarot Masters, ensuring instant guidance anytime, anywhere. It provides various reading types, including compatibility spreads, yes/no tarot, 1-card, 3-card, 6-card, twin flames, relationship, daily transit, weekly transit, and career readings. Tarot Master aims to make spiritual guidance accessible and budget-friendly, offering expert insights without the traditional high costs.

Paraspot AI

Paraspot AI

60%

Paraspot AI revolutionizes property inspections with its AI-powered remote scanning technology, designed to automate the entire inspection process for property managers. The platform allows tenants to perform AI-guided and verified inspections from any mobile device, significantly reducing costs and boosting tenant satisfaction. It generates instant, detailed property reports highlighting damages like cracks, stains, and missing items, which helps eliminate security deposit disputes and ensures faster turnovers. Paraspot AI offers tailored solutions for multifamily, single-family, student housing, furnished rentals, co-living, and short-term rentals, providing a comprehensive portfolio management dashboard to track move-ins, move-outs, and inspection reports in one centralized location. The tool aims to put property operations on autopilot, improving efficiency and accuracy while reducing operational costs by up to 90%.

talk2arxiv

talk2arxiv

60%

talk2arxiv is an open-source Retrieval-Augmented Generation (RAG) system specifically designed for academic paper PDFs. It enables users to chat with any ArXiv paper by simply modifying the paper's URL. The system features PDF parsing using GROBID for efficient text extraction, a custom chunking algorithm that organizes text by logical sections and recursive subdivision, and Cohere's EmbedV3 model for accurate text embeddings. It integrates with Qdrant for vector database storage and querying, which also caches research papers to avoid re-embedding. A reranking process ensures contextual relevance based on user input. The frontend is built with Typescript, ReactJS, TailwindCSS, and NextJS, while the backend utilizes Flask, Gunicorn, and Nginx.

Ema

Ema

60%

Ema is a Universal AI Employee solution designed for enterprises, leveraging sophisticated AI Agents to automate tasks and enhance productivity across all roles and industries. It goes beyond simple automation by learning, adapting, and evolving to meet business needs. Ema offers pre-built AI Agents and a Generative Workflow Engine™ to conversationally activate new AI employees for complex workflows. It is pre-integrated with hundreds of applications, making it easy to configure and deploy. Ema prioritizes data governance, redacting sensitive information before public LLM processing, ensuring compliance with leading standards, top-tier encryption, and customizable private models. Its proprietary EmaFusion™ model, with 2T+ parameters, maximizes accuracy at the lowest cost by intelligently blending public and private models, ensuring future-proof adaptability.

music_recommender

music_recommender

60%

music_recommender is an open-source project designed to provide personalized music recommendations through deep learning. Utilizing Keras and TensorFlow, it processes music data to understand patterns and user preferences, enabling the generation of relevant suggestions. The repository includes various Jupyter notebooks for tasks such as creating feature vectors, concatenating feature arrays, calculating cosine similarity for recommendations, and evaluating holdout sets. It also features Python scripts for classifying music genres and selecting model data. This tool is ideal for developers and data scientists interested in building or experimenting with deep learning-based music recommendation systems.

TileRT

TileRT

60%

TileRT is an open-source, tile-based runtime engineered for ultra-low-latency Large Language Model (LLM) inference. It aims to push the boundaries of LLM latency without compromising model size or quality, allowing models with hundreds of billions of parameters to achieve millisecond-level time per output token (TPOT). Unlike traditional inference systems optimized for high-throughput batch processing, TileRT prioritizes responsiveness, making it ideal for applications like high-frequency trading, interactive AI, real-time decision-making, and AI-assisted coding. It achieves this by decomposing LLM operators into fine-grained tile-level tasks and dynamically rescheduling computation, I/O, and communication across multiple devices to minimize idle time and improve hardware utilization. TileRT currently supports models like GLM-5 and DeepSeek-V3.2 and offers Multi-Token Prediction (MTP) for efficient longer output generation.

table-transformer

table-transformer

60%

Table Transformer (TATR) is a deep learning model developed by Microsoft for extracting tables from unstructured documents, including PDFs and images. Based on object detection, TATR can be trained to work across various document domains, with pre-trained model weights available for the PubTables-1M dataset. The repository also provides the official code for the PubTables-1M dataset, a large-scale dataset for table detection, structure recognition, and functional analysis, and the GriTS evaluation metric for table structure recognition. Researchers and developers can use TATR to detect and recognize tables, convert them to HTML or CSV, and train custom models for specific needs.

wolfcha

wolfcha

60%

Wolfcha is an AI-powered social deduction game, similar to Werewolf or Mafia, where every player is controlled by advanced large language models (LLMs). This innovative game allows users to experience the core appeal of Werewolf—logical deduction, verbal sparring, and reading between the lines—without needing a large group of human players. It features a dual-layer AI roleplay system where virtual players with unique personalities take on Werewolf roles, generating real-time, unpredictable conversations. Wolfcha also serves as an AI model arena, integrating and showcasing various top LLMs like DeepSeek V3.2, Qwen3-235B-A22B, Kimi K2, Gemini 3 Flash, and Seed 1.8 (ByteDance). Players can observe which models reason sharply or seem "adorably clueless," effectively acting as a hidden Turing test. The game offers an immersive retro design style with dynamic interactions like eye-blink transitions and character lip-sync animations.

witsy

witsy

60%

Witsy is an open-source project available on GitHub, functioning as a desktop AI assistant and a universal MCP client. While the GitHub repository itself doesn't provide extensive details on its specific AI capabilities or use cases, its description as an "AI assistant" suggests it aims to help users automate tasks and manage workflows directly from their desktop environment. The mention of a "universal MCP client" indicates potential for integration with various platforms or protocols, making it a versatile tool for developers or technical users looking to customize their AI-driven automation. The project has since moved to a new home under Kochava-Studios.

textgenrnn

textgenrnn

60%

textgenrnn is a Python 3 module built on Keras/TensorFlow designed for creating character-level recurrent neural networks (char-RNNs). It enables users to easily train text-generating neural networks of any size and complexity on any text dataset. The tool incorporates modern neural network architectures, including attention-weighting and skip-embedding, to accelerate training and enhance model quality. Users can train and generate text at either the character or word level, configure RNN size, layer count, and use bidirectional RNNs. It supports training on generic input text files, including large ones, and allows for GPU-trained models to generate text on a CPU. Additionally, textgenrnn offers a powerful CuDNN implementation for faster GPU training and supports contextual labels for improved learning and results.

Blue Prism

Blue Prism

60%

SS&C Blue Prism provides agentic automation solutions for enterprises, specializing in robotic process automation (RPA), business process management (BPM), and artificial intelligence (AI). The platform is designed to handle high-stakes, high-compliance environments across various industries like banking, healthcare, and insurance. It emphasizes built-in governance, proven execution, and a clear path to value, helping businesses operate faster, safer, and smarter. Blue Prism's agentic AI allows agents to make decisions and take actions autonomously, reducing the need for constant human oversight. The platform integrates with various AI tools and offers a Digital Exchange with over 2,000 automation software components, including generative AI and agentic AI.

TheAgentCompany

TheAgentCompany

60%

TheAgentCompany is an open-source benchmark designed to evaluate the performance of LLM agents on consequential, real-world tasks within a simulated software company environment. It allows for assessing how well AI agents can accelerate or autonomously perform work-related tasks by interacting with the web, writing code, running programs, and communicating. The platform offers diverse task roles, data types, and a comprehensive scoring system with multiple evaluation methods, including deterministic and LLM-based evaluators. It features simple one-command operations for environment setup and quick system resets, making it an extensible framework for adding new tasks and evaluators. The benchmark is available on GitHub and supports integration with platforms like OpenHands.

trae-agent

trae-agent

60%

Trae Agent is an LLM-based agent designed for general-purpose software engineering tasks, offering a transparent and modular architecture for researchers and developers. It provides a powerful command-line interface (CLI) that can interpret natural language instructions and execute intricate software engineering workflows using various tools and LLM providers. Key features include Lakeview for concise summarization of agent steps, multi-LLM support for providers like OpenAI, Anthropic, and Google Gemini, and a rich tool ecosystem for file editing, bash execution, and sequential thinking. The agent also offers an interactive mode for iterative development, detailed trajectory recording for debugging, and flexible YAML-based configuration. It is easily installed via pip and supports Docker for isolated task execution.

terminal-bench

terminal-bench

60%

terminal-bench is an open-source benchmark designed to evaluate the performance of AI agents, specifically Large Language Models (LLMs), in realistic terminal environments. It provides a comprehensive suite of tasks that challenge agents with complex, end-to-end scenarios, ranging from compiling code to training models and setting up servers. The tool consists of a dataset of tasks, each with an English instruction, a test script for verification, and a reference solution, along with an execution harness that connects the language model to a sandboxed terminal environment. This setup ensures reproducible and practical evaluation of system-level reasoning. It is currently in beta with approximately 100 tasks, with plans for significant expansion, and welcomes community contributions for new and challenging tasks.

Focus Buddy

Focus Buddy

60%

Focus Buddy is an AI co-pilot designed to enhance productivity by providing AI-powered focus sessions. It actively co-works with users, learning their work patterns, managing to-do lists, and helping to avoid procrastination. The tool offers accountability through AI coach check-ins, assisting users in overcoming barriers like perfectionism and getting started on tasks. It also provides insights into individual work habits, identifying burnout patterns, distractions, and peak productivity times, with weekly reports and upcoming real-time coaching. Focus Buddy aims to be affordable and accessible, offering a free general use version and a personalized paid option.

Refiners IC-Light

Refiners IC-Light

60%

Refiners IC-Light is an AI-powered tool available as a Hugging Face Space that allows users to easily enhance the lighting and appearance of their images. By uploading an image, users gain control over various lighting preferences and settings, enabling them to customize the illumination to their exact needs. The tool then processes these inputs to generate a relighted image, offering a straightforward way to improve visual aesthetics without complex editing software. This makes it accessible for anyone looking to quickly adjust the lighting of their photos for better presentation or artistic effect.

AtmosAi

AtmosAi

60%

Atmos AI is an agentic AI marketing engine designed to autonomously plan, execute, and optimize marketing campaigns for mid-market companies. It features over 180 specialized AI agents that work together across 36 marketing modules, covering areas like content, ads, email, SEO, lead generation, and analytics. Users can access this engine through three distinct brands: Marketing Titan for the full platform, Lead Titan AI for lead intelligence and outreach, and Darwin AI for a natural-language AI Chief of Staff. The platform offers flexible control levels, from full human approval to guided autonomy and full automation, allowing users to set the desired level per campaign or module. Built over 2.5 years, Atmos AI aims to provide a comprehensive solution that runs marketing rather than just recommending actions.

Video-MME

Video-MME

60%

Video-MME is the first-ever comprehensive evaluation benchmark designed to assess the capabilities of Multi-modal Large Language Models (MLLMs) in video analysis. It covers a wide range of visual domains, temporal durations, and data modalities, including short, medium, and long-term videos (from 11 seconds to 1 hour). The benchmark comprises 900 videos totaling 254 hours and 2,700 human-annotated question-answer pairs. It integrates multi-modal inputs beyond video frames, such as subtitles and audios, to provide a full-spectrum evaluation. Video-MME is suitable for both image MLLMs and video MLLMs, offering a robust framework for evaluating model performance in understanding and processing sequential visual data.

streaming-vlm

streaming-vlm

60%

StreamingVLM is an innovative AI tool designed for real-time understanding of effectively infinite video streams. Developed by mit-han-lab, it addresses common challenges in long-video analysis by maintaining a compact KV cache and aligning training directly with streaming inference. This approach efficiently avoids the quadratic cost associated with traditional methods and mitigates the pitfalls of sliding-window techniques. The system is capable of running at up to 8 frames per second (FPS) on a single H100 GPU, offering stable and efficient video processing. It has demonstrated superior performance, winning 66.18% against GPT-4o mini on a new long-video benchmark and also enhances general Video Question Answering (VQA) capabilities without requiring task-specific fine-tuning. The project provides scripts for environment setup, inference, supervised fine-tuning (SFT), and various evaluations including OVOBench and VQA tasks.

streaming-llm

streaming-llm

60%

StreamingLLM is an innovative open-source framework designed to address the challenges of deploying Large Language Models (LLMs) in streaming applications that require processing infinite-length inputs. It introduces the concept of "attention sinks" to efficiently manage Key and Value (KV) states, allowing LLMs to generalize to infinite sequence lengths without fine-tuning. This approach prevents the performance degradation seen in traditional window attention methods when text length exceeds cache size. StreamingLLM enables models like Llama-2, MPT, Falcon, and Pythia to perform stable and efficient language modeling with millions of tokens, offering up to a 22.2x speedup over sliding window recomputation baselines. It is particularly optimized for scenarios such as multi-round dialogues where continuous operation without extensive memory or dependency on past data is crucial.

codeflying

codeflying

60%

CodeFlying is an innovative AI-powered platform designed for "vibe coding," allowing users to build full-stack applications simply by describing their ideas to an AI. This no-code solution streamlines the app development process, enabling the creation of web apps, mobile apps, and even WeChat mini-programs in minutes. It aims to democratize app creation, making it accessible to individuals without extensive coding knowledge. The platform focuses on rapid prototyping and deployment, transforming conversational input into functional applications, marking a new era in app development.