> [!abstract|bg-gray no-i ttl-c c-p-sm] Notes
> 1. Look into [Brave Creator](https://creators.brave.com) program?
> 2. Obsidian Bases
## File Over App
[_File over app_](https://stephango.com/file-over-app) is a philosophy: if you want to create digital artifacts that last, they must be files you can control, in formats that are easy to retrieve and read. Use tools that give you this freedom.
*File over app* is an appeal to tool makers: accept that all software is ephemeral, and give people ownership over their data.
> If you want your writing to still be readable on a computer from the 2060s or 2160s, it's important that your notes can be read on a computer from the 1960s.
You should want the files you create to be durable, not only for posterity, but also for your future self. You never know when you might want to go back to something you created years or decades ago. Don't lock your data into a format you can't retrieve.
Terms and policies are not self-guaranteeing. A company may promise the privacy of your data, but those policies can change at any time. Changes can retroactively affect data you have spent years putting into the tool. Examples: [Google](https://x.com/kepano/status/1682829662370557952), [Zoom](https://x.com/kepano/status/1688606865058574339), [Dropbox](https://x.com/kepano/status/1735032935336829230), [Tumblr](https://x.com/kepano/status/1762864738499952756), [Slack](https://x.com/kepano/status/1791266503456907554), [Adobe](https://x.com/kepano/status/1798459810981220621), [Figma](https://x.com/kepano/status/1808167319694368999).
A self-guaranteeing promise about privacy gives you proof that the tool cannot access your data in the first place.
Encoding values into a governance structure is not self-guaranteeing. Given enough motivation, the corporate structure can be reversed. The structure is not in your hands. Example: [OpenAI](https://en.wikipedia.org/wiki/Removal_of_Sam_Altman_from_OpenAI).
Open source *alone* is not self-guaranteeing. Even open source apps can rely on data that is stuck in databases or in proprietary formats that are difficult to switch away from. Open source is not a reliable safeguard against the biases of [venture capital](https://stephango.com/vcware). Examples: [Omnivore](https://x.com/kepano/status/1851555417165598790), [Skiff](https://news.ycombinator.com/item?id=39396130).
When you choose a tool, the future of that tool is always ambiguous. On a long enough timeline the substrate changes. Your needs change, the underlying operating system changes, the company goes out of business or gets acquired, better options come along.
It is possible to accept the ambiguousness of a tool's future if you choose tools that make self-guaranteeing promises.
## Conserving
Our bias is to always add more. More rules, more process, more code, more features, more stuff. Interdependencies proliferate, and gradually strangle us. Systems want to grow and grow, but without pruning, they collapse. Slowly, then spectacularly.
When a piece of trash drifts across the beach, it is our duty to pick it up so the next person can enjoy a pristine shoreline. When a thousand pieces litter the beach, it is too late. We can only lament the landscape. *That's just how beaches are now*.
A good system is designed to be periodically cleared of cruft. It has a built-in counterbalance. Without this pressure, our bias drives us to add band-aid after band-aid, until the only choice is to destroy the whole system and start from scratch.
Why is it so much easier to add than to remove? Maybe because we attach our identity to what is visible. But there is a difference between the ornamentation that defines our [style](https://stephango.com/style) and the vestigial burdens we carry.
Remember those who did the invisible work of removing. Their legacy was not to build a sand castle, but to care for the beautiful beach on which we play.
## Style
Oscar Wilde once said:
> "Consistency is the last refuge of the unimaginative."
When it comes to ideas, I agree — allow your mind to be changed. When it comes to process, I disagree. Style emerges from consistency, and having a style opens your imagination. Your mind should be flexible, but your process should be repeatable.
Style is a set of constraints that you stick to.
You can explore many types of constraints: colors, shapes, materials, textures, fonts, language, clothing, decor, beliefs, flavors, sounds, scents, rituals. Your style doesn't have to please anyone else. Play by your own rules. Everything you do is open to stylistic interpretation.
A style can be a system, a pattern, a set of personal guidelines. Here are a few of mine:
- I wear monochromatic clothing without logos
- I use `YYYY-MM-DD` dates everywhere
- I pluralize tag and folder names (e.g. `#people` not `#person`)
- I use [plain text files](https://stephango.com/file-over-app) for all my writing
- I ask myself [40 questions every year](https://stephango.com/40-questions)
- I meal prep lunches every week, shave my head twice a week
- I write [concise essays](https://stephango.com/concise), less than 500 words
Collect constraints you enjoy. Unusual constraints make things more fun. You can always change them later. This is *your* style, after all. It's not a life commitment, it's just the way you do things. For now.
Having a style collapses hundreds of future decisions into one, and gives you focus. I always pluralize tags so I never have to wonder what to name new tags.
Style gives you leverage. Every time you reuse your style you save time. A durable style is a great investment.
Style helps you know when you're breaking your constraints. Sometimes you have to. And if you want to edit your constraints, you can. It will be easier to adopt the new constraints if you already had some clearly defined.
You don't need a style for everything. Make a deliberate choice about what needs consistency and what doesn't.
If you stick with your constraints long enough, your style becomes a cohesive and recognizable [point of view](https://stephango.com/in-good-hands).
## In Good hands
There is a feeling I search for: *being in good hands*. It is the feeling I look to give and the feeling I look to receive.
I know I am in good hands when I sense a cohesive point of view expressed with attention to detail.
I can feel it almost instantly. In any medium. Music, film, fashion, architecture, writing, software. At a Japanese restaurant it's what *omakase* aims to be. I leave it up to you, chef.
When I am in good hands I open myself to a state of curiosity and appreciation. I allow myself to suspend preconceived notions. I give you freedom to take me where you want to go. I immerse myself in your worldview and pause judgement.
I want to be convinced of something new. I want my mind to be changed. Later I may disagree, but for now I am letting the experience soak in.
That trust doesn't come easily. As an audience member it's about feeling cared for from the moment I interact with your work. It's about feeling a well-defined point of view permeate what you make.
If my mind was changed, I must have been in good hands.
## Concise explanations accelerate progress
If you want to progress faster, write concise explanations. Explain ideas in simple terms, strongly and clearly, so that they can be rebutted, remixed, reworked — or built upon.
Concise explanations spread faster because they are easier to read and understand. The sooner your idea is understood, the sooner others can build on it.
Concise explanations accelerate decision-making. They help everyone understand the idea and decide whether to agree with it or not.
Concise explanations make ideas useful. One idea can more easily be combined with another idea to form a third idea.
Concise explanations work at every scale. From your own thinking, to the progress of an entire organization, community, or civilization.
Leadership is built on concise explanations. Without concise explanations you have no foundation to build on.
## Don't delegate understanding
There is a parasite, I see it everywhere. It consumes your health and wealth. It preys on ignorance and is easy to catch. It’s so common you may not even notice you have it.
The parasite has a simple and attractive proposition: let me take care of this hard thing for you. Trust me, I know better.
Instead of understanding it yourself, you choose to give the parasite control over your health, education, money, housing, business, identity, data, infrastructure, climate, justice. Even your beliefs.
The parasite has three stages: acceptance, extraction, intervention.
**First is acceptance**. Everyone else seems to have the parasite already. You are expected, even encouraged, to accept the parasite into your life. You are invited to follow the norm, outsource, consume. It’s okay! Use all the services and amenities. Satisfy your desires. Eat the cheap food, watch the cheap media. Your money and time are meant to be spent. Show off what you got in exchange. Please do not try to understand how it works, it’s too complicated for you. The parasite wants you fattened. Literally and figuratively. You are paying the parasite for the privilege of being ripened.
**Second is extraction**. Under the influence of the parasite, you have developed unhealthy habits and you are suffering the consequences. Stress, anxiety, obesity, disease, fear, lethargy, decay. To dampen these problems you pay the parasite for help — support, medicine, loans, fines, rent, taxes. Enforcement of some homeostasis. You try to abate the issues, but you don’t have a stable foundation to build on. You have ignored the root causes. The parasite thrives. You are paying the parasite to be harvested, milked, sucked dry.
**Third is intervention**. The side effects of the parasite’s extraction have reached a critical level. The parasite tells you it’s an emergency. You need doctors, lawyers, firefighters, a military effort. You’re in a surgery room, a court room, a psychiatric ward, a jail cell. The disease can no longer be controlled, it has festered. The flame has turned into a raging fire that needs to be put out. You are paying the parasite to go back to square one.
The three stages of the parasite are interdependent. Every stage benefits someone who is not you. Everyone tells you this is just the way it is. Never mind that the parasite is living large.
Why? Extraction and intervention pay well. Education and prevention do not. The incentives are aligned to make the parasite persuasive. You are alone against a coordinated system that is exceedingly effective at packaging problems you should never have with solutions you should never need. A symbiotic loop.
You must recognize the parasite in its earliest form.
To inoculate yourself don’t delegate understanding. If you build your own understanding you will be the one who earns the dividends.
## Bases
**Using Functions**
Bases includes a range of built-in functions that can be used in filters (to select which notes to include) and formulas (to create new, derived data columns from existing properties). These functions allow for sophisticated data manipulation and querying directly within the Bases UI or its underlying YAML definition.
Here are a few examples of helpful functions:
- **`contains(target, query)`**:
- Checks if a text property or list contains a specific string or item. Useful for filtering notes that mention a keyword or have a specific tag in a list.
- **`if(condition, value_if_true, value_if_false)`**:
- Allows for conditional logic in your formulas. For example, `if(property.status == "done", "Complete", "In Progress")`.
- **`dateAfter(date1, date2)`**:
- Checks if the first date is after the second. Excellent for time-sensitive data, like tasks due after a certain date.
- **`sum(property_name)`**:
- While more advanced aggregation and grouping features are on the roadmap, the table view already supports aggregation like `sum()`. In a view definition, you can specify `agg: "sum(price)"` to calculate the total of a 'price' property for grouped items, for instance. This is a key feature for creating summary dashboards.
### ==.base== File Format
One of Obsidian's core strengths, and a major reason for its dedicated user base, is its commitment to open, user-owned data. Bases continues this tradition:
- **Open Technology Focus**: Obsidian prioritizes open, community-driven technologies like Markdown, YAML, and JSON, avoiding proprietary lock-in.
- **Transparent New Formats**: When a new feature requires a new format, Obsidian ensures it's human-readable and editable. The specification for the `.base` file is simple YAML with a defined schema. You can inspect it, edit it manually if you wish, and understand how your data is structured.
- **Future-Proofing**: Open formats are more likely to be compatible with future tools, including AI agents and other applications.
- **No Licensing Fees**: You own your data and the format it's stored in. You can edit your data in any app in a common text editor.
- **Open Sourcing**: Some formats, like the `.canvas` file format, are even open-sourced, further demonstrating this commitment.
The introduction of a new `.base` file format might initially raise concerns for some, myself included. I learned the hard way with Evernote how frustrating it can be to use a proprietary format. Companies use these formats to lock you into their ecosystem. You'd like to leave but you can't because your data is stuck in their system and is difficult or even impossible to export into another format.
Thankfully, Obsidian puts all of these fears to rest with the new `.base` file format. It's just a simple YAML text file, editable in any text editor, and it uses a simple human readable syntax which is [documented here](https://help.obsidian.md/bases/syntax). This file is used to define things like how you would like to sort or group results. These are many of the same things that we were already defining in Dataview DQL. The good news is that this means it should be quite easy to ask any high quality LLM (like chatGPT) to convert your old Dataview queries into new `.base` files.
Beyond the technical underpinnings, what truly makes Bases exciting is its accessibility.
## API Glossary
[Link](https://brave.com/search/api/glossary/)
A glossary of the concepts behind modern AI search, chatbots, and retrieval systems—from RAG and semantic search to agentic search. Each entry pairs a concise definition with how the term works and where it applies.
### Agentic search
### AI Web Crawler
### Chunking
### Context Window
### Conversational search
### Crawl Budget
### Entity Extraction
[Entity extraction](https://brave.com/search/api/glossary/entity-extraction/) is the natural-language-processing (NLP) task of automatically identifying meaningful things mentioned in text and labeling each one by type. These things can include people, organizations, places, dates, products, and other domain-relevant entities. In many systems, named entity recognition (NER) is a core part of entity extraction, but entity extraction can also be broader depending on the schema and use case.
In short: entity extraction finds names, places, and other entities in text and tags what each one is, turning prose into structured facts that machines can search and use.
##### How Entity Extraction Works
The goal is to move from raw text to a labeled list of entities, usually in a few stages:
1. Preprocessing: The text is cleaned and split into tokens and sentences.
2. Detection: The system locates spans of text that refer to entities. For example, "Brave Software was founded in San Francisco in 2015" could be parsed as Brave Software [ORG], San Francisco [LOC], 2015 [DATE].
3. Classification: Each detected span is assigned a type such as person, organization, location, date, or a custom domain category. Early systems used rules and dictionaries for this step; most production systems today use transformer-based models that infer type from context rather than fixed patterns.
4. Entity linking (optional next step): Extracted entities can then be passed to entity linking, which matches each entry to a unique record in a knowledge base, disambiguating terms based on context.
5. Output: Entities, types, and links are returned as structured data, often attached to the document as metadata or fed into a knowledge graph.
$1
The defining characteristic is that entity extraction does not just find words; it identifies what those words refer to and what kind of thing each one is. It's this latter step that converts text into a machine-usable structure.
##### Entity Extraction vs. Keyword Extraction, Entity Linking, and relation Extraction
- **Keyword extraction** pulls out frequent or salient words and phrases without saying what they are; entity extraction identifies specific real-world things and labels each by type.
- **Entity linking** connects an extracted entity to a unique record in a knowledge base such as [Wikidata](https://www.wikidata.org/), resolving which real-world thing is meant; entity extraction is the prior step that finds and types the mention.
- **Relation extraction** identifies how entities relate to one another, such as "Brave Software makes a search API"; entity extraction identifies the entities those relations connect.
##### Where Entity Extraction is Used
- **Knowledge graphs**: Extracted entities and their types become nodes that populate a knowledge graph.
- **Search and retrieval**: Tagging documents with entities improves filtering, faceting, and matching entities in a query to content, which helps both classic search and AI answer engines.
- **Content enrichment**: Articles, support tickets, and product pages can be tagged automatically with the entities they mention, creating richer metadata.
- **Structured data and markup**: Extracted entities can map to [schema.org](https://schema.org/) types, helping crawlers and AI systems understand what a page is about.
- **Question answering**: Identifying entities in a question and in candidate answers helps a system retrieve and verify the right facts.
Common tools include open-source libraries such as [spaCy](https://spacy.io/) and models hosted on [Hugging Face](https://huggingface.co/), alongside managed cloud APIs. Entity extraction is most useful when a system needs to move from string matching to understanding real-world concepts.
##### Related terms
Named entity recognition (NER), entity linking, relation extraction, [knowledge graph](https://brave.com/search/api/glossary/knowledge-graph/), [semantic search](https://brave.com/search/api/glossary/semantic-search/), [embeddings](https://brave.com/search/api/glossary/vector-embeddings/), structured data, natural language processing (NLP), schema.org.
### Function Calling
[Function calling](https://brave.com/search/api/glossary/function-calling/) is a capability that lets a large language model (LLM) request the use of an external tool, API, or piece of code by producing a structured, machine-readable call instead of a plain-text reply.
The model does not run the code itself; it decides which function to invoke and with what arguments, while the application then executes that function, and the result is handed back to the model to incorporate into its final answer. Function calling is often used interchangeably with "tool use" or tool calling, though strictly speaking it refers to developer-defined functions, a subset of the broader tool-calling pattern.
### Hallucination
A [hallucination](https://brave.com/search/api/glossary/hallucination/) is output from an AI model (especially a large language model) that is false, fabricated, or ungrounded, yet presented fluently and confidently as if it were true. ("Ungrounded" in this case meaning the output is unsupported by the model's training data or the sources it was given.) Hallucinations range from invented citations and fake statistics to plausible-sounding details that no source actually supports, and can also arise when a model misreads, misattributes, or overgeneralizes from stale or partial sources.
### Hybrid search
[Hybrid search](https://brave.com/search/api/glossary/hybrid-search/) is a retrieval approach that combines keyword (lexical) search and semantic (vector) search as two complementary methods, with results then merged into a single ranked list. By running both at once, hybrid search combines exact-term matches (which keyword search is good at) with meaning-based matches (which semantic search is good at). In mixed-query workloads, this often improves retrieval robustness compared with using either method alone.
### Intent Classification
[Intent classification](https://brave.com/search/api/glossary/intent-classification/) is the natural-language-processing task of determining what a user is trying to accomplish from their input, by assigning a query or message to one of a set of predefined intent categories. It answers the question "what does this person want?" (for example, whether a search query is informational, navigational, or transactional, or whether a chatbot message is a greeting, a complaint, or a request to make a reservation).
### Knowledge Cutoff
A knowledge cutoff is the point in time after which a language model has no training data, and therefore no built-in awareness of events, facts, or changes that occurred later. Because a model's knowledge is effectively frozen at training time, anything that happened after its cutoff is not part of its built-in knowledge unless newer information is supplied at runtime through external context, retrieval, or tools.
### Knowledge Graph
A knowledge graph is a structured representation of knowledge as a network of entities (e.g. people, places, products, or concepts) and the relationships that connect them. Instead of storing information as isolated text, a knowledge graph records facts as nodes (the entities) linked by edges (the relationships), often with types and attributes attached. This allows both machines and people to more easily understand how things relate to one another.
### Large Language Model (LLM)
A large language model (LLM) is an artificial-intelligence model trained on vast amounts of text to understand and generate human language. It learns the statistical patterns of language so it can predict and produce the next piece of text in a sequence, which (at sufficient scale) lets it answer questions, summarize, translate, write code, and hold a conversation.
### Latency
Latency is the delay between a request and the response to it—basically how long a user or system waits for results. In AI search and language-model applications, it usually means the time from sending a query to receiving an answer, and it is a core measure of how responsive and usable a system feels.
### Programmable search Engine
A programmable search engine is a search engine that developers configure and query through code rather than only through a consumer website. This allows developers to build search directly into their own applications, control how it behaves, and receive results in a structured, machine-readable form. The term "programmable search engine" more broadly describes any search engine exposed for programmatic use, typically through a search API.
### Rate Limiting
Rate limiting is the practice of capping how many requests a client (often an app on a computer or mobile device) can make to a service within a set period of time. When a client exceeds the limit, the service rejects or delays further requests (usually returning an HTTP 429 Too Many Requests response) until the window resets. Rate limiting protects a service from overload, ensures fair access among users, and enforces the usage tiers that many APIs sell.
### Reranking
Reranking is a second-stage process that reorders an initial list of search results to put the most relevant ones first. A fast first-stage retriever pulls a broad set of candidate results, and the reranker then scores each candidate more carefully against the query (using a slower but more accurate model) and sorts them by that score. Reranking is the precision step that comes after the recall step.
### Retrieval Score
A retrieval score is a number that a search or retrieval system assigns to each candidate result, expressing how relevant or similar the result is to the original user query. The system ranks results by this score, putting the highest-scoring items first, and can use it to decide which results are strong enough to keep. Depending on the method, the score may measure keyword overlap, semantic similarity, or a blend of both.
### Retrieval-augmented Generation (RAG)
Retrieval-augmented generation (RAG) is a technique that improves a language model's answers by retrieving relevant information from an external source, and supplying this information to the model as context before it generates a response. Rather than relying exclusively on what it learned in training, the model answers using documents fetched at query time, so its output can be current, specific, and grounded in real sources.
### Scraper
A scraper is a program that automatically extracts data from websites (including search engine results pages), usually by downloading their pages and parsing the HTML written for human browsers; scrapers can turn content meant for human readability into structured data a computer program can store and reuse. Scrapers have many legitimate uses (including price comparisons, research, archiving, and system monitoring), but depending on what is scraped and how, they can raise legal, technical, and policy issues. In the search context, a scraper pulls results by loading and parsing a search engine's results pages—an approach that many major providers restrict or prohibit through terms of service and anti-automation policies.
### Search Engine Results Page (SERP)
A search engine results page (SERP) is the page a search engine returns in response to a query. On modern search engines, it is usually a blend of elements rather than a simple list: ranked organic (unpaid) links, paid ads, and SERP features such as featured snippets or knowledge panels that answer parts of the query directly. On some queries and engines, AI-generated overviews may also appear, often near the top of the page.
### Semantic search
Semantic search is an approach to search that matches results by meaning and intent, not just exact keyword overlap. In modern systems, this is commonly implemented by representing queries and content as embeddings (numeric vectors that capture semantic relationships), then retrieving items with the most similar vectors. That lets a query like "best wine for seafood" match content such as "good with fish," even when exact terms differ.
### SERP Features
SERP features are enriched modules on a search engine results page that go beyond a standard list of organic links. Common examples include featured snippets, knowledge panels, "People Also Ask" boxes, image and video carousels, local map packs, top-stories blocks, shopping modules, and AI-generated overviews.
### Source Attribution
Source attribution is the practice of identifying and crediting the original sources behind a piece of information, so a reader can see where a claim originated. In AI search and answer engines, it means pairing generated answers with citations (links or references back to the specific pages and passages the answer draws on) rather than presenting conclusions with no traceable origin.
### Structured Output
Structured output is output from a language model that follows a specific, predefined format (most often JSON matching a defined schema) instead of free-form prose. By constraining response shape, structured output makes model output easier for software to parse and act on without brittle cleanup logic.
### Tokenization
Tokenization is the process of breaking text into smaller units called tokens, the pieces a system actually processes. For a language model, tokens are usually sub-word fragments that the model reads as numbers rather than whole words. In classic search, tokenization instead splits text into the terms used to build and query an index. Either way, it is the step that turns raw text into units a system can work with.
### Tool Calling
Tool calling is the capability that lets a large language model (LLM) use external tools like APIs, functions, databases, or services. By producing a structured request to invoke a tool, an application or the model provider can then execute a task and return a result to the model. This is how a model moves beyond generating text to fetching live data and taking actions.
### Vector Database
A vector database is a database built to store and search embeddings—the high-dimensional numeric vectors that represent the meaning of text, images, or other data. Instead of matching records by exact values like a traditional database, a vector database finds the records whose vectors are most similar to a query, which is what makes meaning-based (semantic) search possible at scale.
### Vector Embeddings
A vector embedding (or just an embedding) is a list of numbers that represents the meaning of a piece of data (such as text, image, or audio) as a point in a high-dimensional space. An embedding model produces these vectors so that items with similar meaning land close together, which lets software compare meaning by measuring distance between vectors rather than matching words.
### Web Grounding
Web grounding is the practice of basing a language model's answers on live web content retrieved at query time, so its responses reflect current, verifiable sources rather than only what it learned in training. A grounded model searches the Web for a query, pulls in the relevant pages, and generates its answer from that retrieved content—often citing the sources it used.
### Web search API
A web search API is a service that lets software query a search engine's index of the Web and receive the results as structured data, rather than through a browser or by scraping a search results page. A program issues a request containing a query and parameters and gets back machine-readable results (typically URLs, titles, and text snippets, alongside news, images, videos, and richer content) ready to use in an AI application.
### Zero-click search
Zero-click search is a search session that ends without the user clicking an outbound result. This can happen because the results page provides enough information directly, because the user abandons the query, or because they reformulate and search again.
## Privacy Glossary
[Link](https://brave.com/glossary/)
%%Footer Starts Here%%
---
![[Brain Icon 1.png|center]]
<b><font color="#ffffff"> <center>You might not have noticed it… but your brain did.</center> </font></b>
---
### Tags
### Linked Pages & Footnotes