Audio to MIDI
- Input
- Solo-piano audio
- Output
- Timed symbolic notes
Onsets and Frames is transcription. It does not create a finished vocal recording.
Sources: Onsets and Frames: Dual-Objective Piano Transcription ↗
INTERN PROGRAM / CASE STUDY / OCTOBER 5, 2026 / EDITORIAL EDITION
How an early music-AI research program can become a practical learning brief for a new generation of builders.
Dated editorial edition · Sources checked October 5, 2026. See each source and claim boundary below.
Adventure Corporation archive · Collective research and development contributions alongside Michael Hoydich’s program and product framing. GPT-3 exploration is documented; direct GPT-2 use remains unverified.
How an early music-AI research program can become a practical learning brief for a new generation of builders.
In 2020, Adventure Corporation’s music work asked what would happen when making, listening and collaborating became parts of the same experience. Research boards explored AI music generation, mood classification, audio recognition, browser-based tools and interactive listening. A short design assignment asked how someone might design an Instagram for music. Nearby notes imagined people creating together through dance, drums, singing, visual elements and emerging interfaces.
For an internship program, that was a useful kind of question. It crossed computer science, product design, media and business without requiring every participant to solve the entire system. One person could compare models. Another could map audio formats. Others could prototype an interaction, build a player component or explain what the group had learned. The shared challenge gave small assignments a larger purpose.
The surviving record supports a case study in research, product framing and learning through making. It contains completed research tasks and concrete design proposals. It does not establish that the team trained a new music foundation model, launched the imagined AI soundtrack service, or achieved a particular audience or revenue result. That boundary is part of the lesson: an interesting direction becomes credible when each claim has an inspectable artifact.
This work belongs to the period before ChatGPT’s public launch. GPT-2 was announced in February 2019, and its largest public release followed that November. OpenAI introduced MuseNet in April 2019 and Jukebox in April 2020. The GPT-3 paper appeared in May 2020; the OpenAI API entered private beta that June. ChatGPT arrived in November 2022, initially based on the GPT-3.5 series.
Those names describe different systems and interfaces. MuseNet worked with symbolic musical sequences. Jukebox generated raw audio. GPT-3 provided a general text interface through the API. A music project could investigate all three without using a single interchangeable “AI engine.” The distinction matters when reconstructing what a team actually did.
The archive explicitly names GPT-3. Michael Hoydich’s signed July 2020 essay connects the new API with Jukebox and Magenta, and a later signed account says the team had GPT-3 beta access and had begun exploring avatars, agents and interactivity. Direct GPT-2 use has not been verified. MuseNet’s official description links its underlying approach to GPT-2, which offers one possible explanation for a remembered GPT-2 connection, but that remains an inference.
Public technology sources: GPT-2 announcement ↗, GPT-2 full release ↗, MuseNet ↗, Jukebox ↗, GPT-3 paper ↗, OpenAI API announcement ↗, Introducing ChatGPT ↗.
A particularly useful piece of evidence is a team-status snapshot created and last updated on June 27, 2020. It records a music-AI research assignment as completed on June 26. The brief asked for a Whimsical overview of projects, tools and experiments, with Jukebox and Magenta supplied as starting points. The same snapshot assigns separate work on audio file types and music recommendations. Research was divided into understandable pieces that could be shared back with the group.
The Music AI board contains a landscape of generation, transcription and supporting infrastructure. It discusses Jukebox, AIVA, Magenta Studio, Onsets and Frames, TensorFlow and cloud machine-learning services. It also records technical questions such as long input sequences and codebook collapse. Several passages summarize external research, so first-person descriptions of training on a large dataset must be attributed to the original model researchers, rather than read as the team’s own training history.
This makes the board valuable as a learning artifact. It connects unfamiliar vocabulary to possible applications, while giving peers a place to compare alternatives. Its strongest demonstrated result is a completed, reusable research overview. A board entry alone cannot demonstrate an installation, a successful model run or a shipped feature.
The contribution record also matters. File metadata lists Mike as the creator of many boards, while the task snapshots name participant assignments. A fair account credits Mike’s program and product framing alongside the participants’ research and development contributions. Board ownership is not a substitute for an authorship record.
The Soundtrack AI Research board organizes its thinking around datasets, AI and product. It asks whether labels could help connect a desired mood with music, including the possibility of interpreting an image as input. It links AudioSet, MusicBrainz, AIVA and MuseNet, and raises licensing as part of the investigation. These are proposals and research notes, rather than a verified end-to-end implementation.
A related AIME board makes the customer situation more tangible: a person editing a video wants a soundtrack and supplies parameters. It also explores recommendation, beat patterns, collaborative music creation and the ancestry of a work as it is reused. Questions about ownership and contribution appear inside the product conversation itself.
The value of this progression is practical. “Explore AI music” can expand forever. “Help a video editor find a suitable soundtrack for a short scene” gives a team something to test. It suggests an input, a result, a review step and a real user judgment. Generation is one possible part of that experience; search, tagging, editing and attribution may be equally important.
The mood-classification board adds another constraint: audio and lyrics provide different kinds of information, and lyrics are not always available. Its idea of pairing a music mood with a GIF is a clear interaction hypothesis. Today’s learner can reuse the question while evaluating it afresh, rather than inheriting a historical accuracy claim or assuming that a mood label is objective.
The archive’s 2021 AI NFT Creation board preserves a useful reminder that capability and usability advance at different speeds. It records Jukebox limitations around sampling time, audible quality and longer musical structure. OpenAI’s original release described roughly nine hours to render one minute of audio. That historical limitation would have mattered enormously to an interactive product, regardless of how exciting the samples sounded.
The answer is to match the experiment to the constraint. An offline listening study, a curated comparison or a queued generation workflow asks something different from a live instrument. A transcription tool asks another question again: Magenta’s Onsets and Frames converts solo-piano audio to MIDI, rather than generating a finished vocal track.
A separate research board proposes spiking-neural-network audio experiments, with stages for data preparation, encoding, classification, validation and a web interface. It includes examples from related prior experiments, but the reviewed materials do not verify completion of the proposed audio system. The useful lesson is the staged plan and its uncertainties. Claims of novel research or performance gains would require code, data, baselines and reproducible results.
Sources: Jukebox ↗, Onsets and Frames ↗.
For a renewed University of El Segundo or IndustryNext program, the strongest inheritance is a repeatable learning process. Begin with a question and a small piece of evidence. Make something that another person can inspect. Ask for critique. Revise. Then leave a short explanation that helps the next participant begin with more context.
A current cohort could study a rights-cleared set of short audio clips, compare human mood annotations, and build a simple soundtrack-selection prototype. The educational outcome would be visible even if generation were never added. Participants would practice defining a dataset, identifying disagreement, designing controls, testing a workflow and communicating the limits of a result.
Roles can remain distinct while the work stays collaborative. Researchers explain the model and its constraints. Designers make uncertainty understandable. Engineers build the smallest reliable interaction. Media contributors document the process. Business participants test whether the proposed result solves a meaningful problem and what would have to be true for it to be sustainable.
Each team should finish with a compact evidence pack: the original question, tool and model versions, rights and source notes, a working artifact or clearly labelled mockup, observations from testing, a contribution record and the next unresolved question. Assessment should reward the quality of this reasoning and the usefulness of the handover, alongside craft.
Later materials extend the creative-AI thread. A board created in December 2022 names ChatGPT, coding assistants and image tools in an AI-assisted team concept. A January 2023 image-generation board combines images with a proposed editorial schedule and magazine copy. These artifacts show an interest in turning emerging tools into creative and organizational workflows.
By 2026, concept boards explicitly separate roles across human direction and several AI tools. They are useful examples of workflow design, but their forecasts and planned launches should remain labelled as projections. Images also need provenance: a plausible-looking concept render does not establish a manufactured product, a commissioned campaign or a completed sale.
The opportunity for PointCast is to make the learning legible. An episode or case-study page can show the original question, a reconstructed research map, a present-day experiment and an honest account of what changed. The audience should be able to distinguish archival evidence, new interpretation and proposed work at a glance.
Early attention to a technology creates possibilities. Turning it into value requires a narrower problem, usable evidence, clear responsibilities and a testable next step. This archive is strongest when it shows how those pieces were being assembled: a shared research task, a product question, a technical limitation and a path for participants to contribute.
Learn, build, ship remains a useful invitation when each verb has a visible outcome. Learn enough to make a justified choice. Build enough to expose the next uncertainty. Ship an artifact and an explanation that someone else can use. The result is a case study that future participants can continue, rather than a finished success story they are expected simply to admire.
Based on a read-only review of selected Adventure Corporation Whimsical materials and public primary research. The reviewed materials distinguish completed research, proposed products and reported exploration. Some boards were edited years after creation; their creation timestamps do not date every item. Private source boards and participant materials are not reproduced. Public research links establish technology history, not independent verification of the team’s use.
University of El Segundo and IndustryNext are presented here as proposed learning/program contexts. This case study makes no claim of accreditation, live enrollment, employment terms or guaranteed commercial results.
Use GPT-3 for the team’s explicitly documented language-model references. GPT-2 usage remains unverified. Do not use “ChatGPT 2” for the 2020 work.
Public releases establish technology history. Archive dates describe recorded milestones or file creation, with later edits explicitly noted. They do not date every idea or demonstrate a shipped product.
19 of 19 cards
Automatic polyphonic solo-piano transcription research.
Sources: Onsets and Frames: Dual-Objective Piano Transcription ↗
First collection of MIDI creativity tools announced.
Sources: Magenta Studio announcement ↗
Language model announcement and initial smaller-model release.
Sources: GPT-2 announcement ↗
Music generation system; the official description relates its general approach to GPT-2.
Sources: MuseNet ↗
Largest 1.5B model released. Team use is unverified.
Sources: GPT-2 full 1.5B release ↗
Original arXiv submission; subsequent revisions are separately dated.
Sources: Language Models are Few-Shot Learners, original submission ↗
Announcement names GPT-3-family weights.
Sources: OpenAI API announcement ↗
June 27 snapshot records completion on June 26; the brief names Jukebox and Magenta. This is a research-deliverable milestone.
Sources:
Signed essay explicitly connects GPT-3, Jukebox and Magenta. The surviving board was created July 16 and edited later.
Sources:
Current content covers labels, datasets and a soundtrack-product hypothesis; later edits prevent dating every node to creation.
Sources:
Current content explores parameter-driven soundtracks, collaborative music and attribution; later edited.
Sources:
Surviving text reports GPT-3 beta access and early avatar/agent exploration; later edited.
Sources:
Research includes generative media and model limitations; not evidence of NFT sales or deployed generation.
Sources:
Initially fine-tuned from GPT-3.5. Do not call the 2020 work “ChatGPT 2.”
Sources: Introducing ChatGPT ↗
OSFD Super Builders board created; surviving copy names ChatGPT and creative/coding tools. Later edited.
Sources:
Image Generation AI, Desktop board created; last edited January 11. Per-image model and prompt provenance are unverified.
Sources:
ICONIC MINIS board explicitly assigns roles to human direction and AI tools. Business figures are projections.
Sources:
University of El Segundo planning board connects the historical learning loop to new studio briefs. It is a proposal, not live enrollment or accreditation.
Sources:
No milestones match. Reset filters to restore both timelines.
Onsets and Frames is transcription. It does not create a finished vocal recording.
Sources: Onsets and Frames: Dual-Objective Piano Transcription ↗
MuseNet works with symbolic sequences; its official history does not establish this team’s use.
Sources: MuseNet ↗
Jukebox researches raw-audio music generation. Its historical latency is a product constraint.
Sources: Jukebox ↗
Original educational reconstruction, prepared October 2026. No archival audio, copied research diagram or historical product screenshot is presented.
Proposed learning briefs, with deliverables that another participant can inspect. Mentoring, hours, pay and eligibility are unresolved; these are not open employment listings.
Sort ten supplied, sanitized claims into research, proposal, implemented artifact, reported result and independently verified result.
Use five self-created or appropriately licensed instrumental clips. Ask three adult volunteers to label mood independently; record disagreement without collecting unnecessary personal data.
Build or paper-prototype a tool for choosing among the rights-cleared clips using scene, energy, duration and instrumentation. Make selection criteria visible.
Compare a transcription task, a symbolic-composition task and a raw-audio-generation task using the cited historical research. If running a current tool, verify its availability, terms and hardware needs first.
Produce a three-minute walkthrough of the question, attempt, observation and next step. Use only publication-approved artifacts.
Six business lessons drawn from research and product framing. These are interpretations, not claims of historical revenue, deployment or commercial success.
A specific video-editing moment creates a more useful experiment than a general promise to transform music.
A shared, source-checked comparison can save a team from building the wrong thing. Its outcome should still be labelled research.
Latency, controllability, quality, cost and rights determine which experience is viable.
Maintain role and artifact credits. A file owner or AI-tool label does not explain who researched, designed, wrote or reviewed the work.
Concept images and a financial model support discussion; user testing, deployment and sales each require their own evidence.
A small artifact with clear sources and a good handover is more teachable than an expansive plan with no demonstrated next step.
Fresh educational reconstruction of the learning loop, not an original 2020 diagram. UES supplies the learning lens; the intern program supplies role and handover questions; IndustryNext supplies customer and sustainability questions.
Sources were checked October 5, 2026. Each card links to the evidence behind its facts; analysis is PointCast editorial interpretation.
MusicVAE and NSynth are supplemental research links. Direct historical use of those models by this team is not established. Public links establish release history, not independent proof of the team’s tool use.